{
  "id": 51142,
  "title": "Welcome",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51142",
  "author_name": "inversion",
  "post_date": "2018-03-05T20:08:12.813000",
  "votes": 27,
  "comment_count": 65,
  "views": 0,
  "content": "<p>Welcome to the TalkingData AdTracking Fraud Detection Challenge! This is a challenge for those who love working on large, imbalanced, binary classification problems. (Or for those who want to build their skills in that area!)</p>\n\n<p>Note that the training data is ~8 Gb uncompressed. For those who want a quick preview, we've added a training sample file that contains 100k random rows. </p>\n\n<p>Please feel free to ask your questions in this thread.</p>\n\n<p>Good luck!</p>",
  "messages": [
    {
      "id": 291215,
      "postDate": "2018-03-05T20:08:12.813Z",
      "content": "<p>Welcome to the TalkingData AdTracking Fraud Detection Challenge! This is a challenge for those who love working on large, imbalanced, binary classification problems. (Or for those who want to build their skills in that area!)</p>\n\n<p>Note that the training data is ~8 Gb uncompressed. For those who want a quick preview, we've added a training sample file that contains 100k random rows. </p>\n\n<p>Please feel free to ask your questions in this thread.</p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "Welcome to the TalkingData AdTracking Fraud Detection Challenge! This is a challenge for those who love working on large, imbalanced, binary classification problems. (Or for those who want to build their skills in that area!)\n\nNote that the training data is ~8 Gb uncompressed. For those who want a quick preview, we've added a training sample file that contains 100k random rows. \n\nPlease feel free to ask your questions in this thread.\n\nGood luck!",
      "votes": 27
    },
    {
      "id": 299949,
      "postDate": "2018-03-21T06:27:43.990Z",
      "content": "<p>There seem to be multiple questions reg IP addresses([1], [2]). @inversion would be great if you could give your inputs here. The main questions seem to be regarding the nature of the function used to anonymize the IP addresses:</p>\n\n<p>Questions are:</p>\n\n<p>(1) whether the anonymization preserves IP range similarity. i.e., if IP1 and IP2 belonged to the same class C subnet originally, will they have similar values after anonymization. (for e.g., a hash-based anonymization will not preserve this IP range similarity)</p>\n\n<p>(2) Is the same anonymization scheme used for both training and test set ? If the same IP address IP1 appears in both training and test sets before anonymization, will it have the same value in both training and test sets after anonymization ?</p>\n\n<p>[1] <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal</a></p>\n\n<p>[2] <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374</a></p>",
      "rawMarkdown": "There seem to be multiple questions reg IP addresses([1], [2]). @inversion would be great if you could give your inputs here. The main questions seem to be regarding the nature of the function used to anonymize the IP addresses:\n\nQuestions are:\n\n(1) whether the anonymization preserves IP range similarity. i.e., if IP1 and IP2 belonged to the same class C subnet originally, will they have similar values after anonymization. (for e.g., a hash-based anonymization will not preserve this IP range similarity)\n\n(2) Is the same anonymization scheme used for both training and test set ? If the same IP address IP1 appears in both training and test sets before anonymization, will it have the same value in both training and test sets after anonymization ?\n \n[1] https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\n\n[2] https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374",
      "votes": 9,
      "replies": [
        {
          "id": 303477,
          "postDate": "2018-03-26T09:04:04.947Z",
          "content": "<p>Same Question</p>",
          "rawMarkdown": "Same Question"
        }
      ]
    },
    {
      "id": 291895,
      "postDate": "2018-03-07T03:30:38.857Z",
      "content": "<p>Can we get more decimal places please?</p>",
      "rawMarkdown": "Can we get more decimal places please?",
      "votes": 7,
      "replies": [
        {
          "id": 292249,
          "postDate": "2018-03-07T17:49:17.087Z",
          "content": "<p>I second this. The leaderboard is somewhat pointless right now.</p>",
          "rawMarkdown": "I second this. The leaderboard is somewhat pointless right now.",
          "votes": 1
        },
        {
          "id": 292317,
          "postDate": "2018-03-07T20:02:13.543Z",
          "content": "<p>Upvote for more decimals</p>",
          "rawMarkdown": "Upvote for more decimals",
          "votes": 2
        },
        {
          "id": 292403,
          "postDate": "2018-03-07T22:28:01.313Z",
          "content": "<p>Cannot agree more.</p>",
          "rawMarkdown": "Cannot agree more."
        },
        {
          "id": 293297,
          "postDate": "2018-03-09T15:45:24.043Z",
          "content": "<p>LB now has 4 sig figs.</p>",
          "rawMarkdown": "LB now has 4 sig figs.",
          "votes": 4
        }
      ]
    },
    {
      "id": 324997,
      "postDate": "2018-05-08T03:01:28.063Z",
      "content": "<p>@ inversion </p>\n\n<p>I guess the Leader board got hammered. \nLet's take a look at this <a href=\"https://www.kaggle.com/umeshsati54/merge-and-avg/data\">BLENDING OF PUBLIC CSV FILES </a></p>\n\n<p>Someone has just downloaded the outputs of several files and used the same files again. I believe these same set of files must have been used by a large number of people ( Some of the best people in the competition are nowhere to be seen )</p>\n\n<p>Is that not PLAGIARISM?????? I understand the importance of Kaggle staying open sourced but at the same time KAGGLE MUST PENALIZE plagiarism as well. </p>\n\n<p>I was so sure that this was outside the rules that I did not use the file. I didn't use that file because I thought I would get disqualified. But what I got in return was getting pushed down from top 7% to 22 % in the Leader board. </p>\n\n<p>Hoping Kaggle can do something about the final results.</p>",
      "rawMarkdown": "@ inversion \n\nI guess the Leader board got hammered. \nLet's take a look at this [BLENDING OF PUBLIC CSV FILES ][1]\n\nSomeone has just downloaded the outputs of several files and used the same files again. I believe these same set of files must have been used by a large number of people ( Some of the best people in the competition are nowhere to be seen )\n\nIs that not PLAGIARISM?????? I understand the importance of Kaggle staying open sourced but at the same time KAGGLE MUST PENALIZE plagiarism as well. \n\nI was so sure that this was outside the rules that I did not use the file. I didn't use that file because I thought I would get disqualified. But what I got in return was getting pushed down from top 7% to 22 % in the Leader board. \n\nHoping Kaggle can do something about the final results.\n\n  [1]: https://www.kaggle.com/umeshsati54/merge-and-avg/data",
      "votes": 4,
      "replies": [
        {
          "id": 325302,
          "postDate": "2018-05-08T09:36:49.770Z",
          "content": "<p>totally agree with you! think that the very first kernel should be published and all others with 100% correlation should be threated as dublicates. So you can improve your solution based on features of this kernel and NOT to copy it blindly.</p>",
          "rawMarkdown": "totally agree with you! think that the very first kernel should be published and all others with 100% correlation should be threated as dublicates. So you can improve your solution based on features of this kernel and NOT to copy it blindly."
        }
      ]
    },
    {
      "id": 291255,
      "postDate": "2018-03-05T21:25:45.887Z",
      "content": "<p>Thanks, @inversion! </p>\n\n<p>I wonder if it is possible for Kaggle to provide larger memory for kernels. I found it is hard to load all the data while the sampled training set is to small to fit model.</p>",
      "rawMarkdown": "Thanks, @inversion! \n\nI wonder if it is possible for Kaggle to provide larger memory for kernels. I found it is hard to load all the data while the sampled training set is to small to fit model.",
      "votes": 3,
      "replies": [
        {
          "id": 291267,
          "postDate": "2018-03-05T21:50:59.443Z",
          "content": "<p>Same question!</p>",
          "rawMarkdown": "Same question!",
          "votes": 1
        },
        {
          "id": 291272,
          "postDate": "2018-03-05T21:55:32Z",
          "content": "<p>Found it is ok to load 10m row but the total is ~ 185m rows  :(</p>",
          "rawMarkdown": "Found it is ok to load 10m row but the total is ~ 185m rows  :("
        },
        {
          "id": 291382,
          "postDate": "2018-03-06T03:08:51.997Z",
          "content": "<p>My plan is to use data stacking and kernel chaining to beat this competition!</p>",
          "rawMarkdown": "My plan is to use data stacking and kernel chaining to beat this competition!",
          "votes": 1
        },
        {
          "id": 291497,
          "postDate": "2018-03-06T08:35:25Z",
          "content": "<p>@Muhammad Alfiansyah, sounds like a great plan. I am working on fit_generator from Keras so I can load/fit the training set batch by batch.</p>",
          "rawMarkdown": "@Muhammad Alfiansyah, sounds like a great plan. I am working on fit_generator from Keras so I can load/fit the training set batch by batch.",
          "votes": 1
        },
        {
          "id": 291542,
          "postDate": "2018-03-06T10:22:07.473Z",
          "content": "<p>@Shujian.M.Coder, looking forward to learn from and working with you!</p>",
          "rawMarkdown": "@Shujian.M.Coder, looking forward to learn from and working with you!",
          "votes": 1
        },
        {
          "id": 291671,
          "postDate": "2018-03-06T16:15:30.093Z",
          "content": "<p>Unfortunately, we aren't able to increase memory for Kernels in the short term.</p>",
          "rawMarkdown": "Unfortunately, we aren't able to increase memory for Kernels in the short term.",
          "votes": 5
        },
        {
          "id": 295862,
          "postDate": "2018-03-14T10:28:56.673Z",
          "content": "<p>I see that some people use AWS/GCP for additional compute capacity. Perhaps you could link to that in your Kernel? These services often provide starter credits and offer highly capable analysis platforms if you have the time to tinker with them. My guess is that Kaggle has a significantly lower budget for this than corporations like Amazon and Google, so unless we want to make Kaggle into a pay site we will have to work within the limits of digital sustainability.</p>",
          "rawMarkdown": "I see that some people use AWS/GCP for additional compute capacity. Perhaps you could link to that in your Kernel? These services often provide starter credits and offer highly capable analysis platforms if you have the time to tinker with them. My guess is that Kaggle has a significantly lower budget for this than corporations like Amazon and Google, so unless we want to make Kaggle into a pay site we will have to work within the limits of digital sustainability."
        },
        {
          "id": 295896,
          "postDate": "2018-03-14T11:37:13.877Z",
          "content": "<p>Thanks, Marissa. I have almost burnt all the free credits from AWS and GCP. That's why I hope I can just run some experiments in Kaggle kernels instead pay by my own.</p>",
          "rawMarkdown": "Thanks, Marissa. I have almost burnt all the free credits from AWS and GCP. That's why I hope I can just run some experiments in Kaggle kernels instead pay by my own.",
          "votes": 1
        }
      ]
    },
    {
      "id": 326113,
      "postDate": "2018-05-09T09:57:40.280Z",
      "content": "<p>@ inversion </p>\n\n<p>Is the Kaggle team still looking into the whole issue of shared CSV's in the Talking data competition? Are you having specific discussions internally on re-evaluating Talking Data based on the duplicate CSVs?</p>\n\n<p>OR is the competition considered to be CLOSED and no further scrutiny is to be done by Kaggle?</p>\n\n<p>Regards\nShanth</p>",
      "rawMarkdown": "@ inversion \n\nIs the Kaggle team still looking into the whole issue of shared CSV's in the Talking data competition? Are you having specific discussions internally on re-evaluating Talking Data based on the duplicate CSVs?\n\n OR is the competition considered to be CLOSED and no further scrutiny is to be done by Kaggle?\n\nRegards\nShanth",
      "votes": 1
    },
    {
      "id": 302624,
      "postDate": "2018-03-24T13:18:30.567Z",
      "content": "<p>Why doesn't kaggle introduce joint ranking for the same scores? Doesn't seem fair that so many people can be eranked lower or higher for same score. It's the difference between winning a medal and not having one.</p>",
      "rawMarkdown": "Why doesn't kaggle introduce joint ranking for the same scores? Doesn't seem fair that so many people can be eranked lower or higher for same score. It's the difference between winning a medal and not having one.",
      "votes": 1,
      "replies": [
        {
          "id": 317665,
          "postDate": "2018-04-22T07:20:17.207Z",
          "content": "<p>the one achieved the score earlier has a higher rank, it is fair. and it is just based on public set anyway</p>",
          "rawMarkdown": "the one achieved the score earlier has a higher rank, it is fair. and it is just based on public set anyway"
        }
      ]
    },
    {
      "id": 291768,
      "postDate": "2018-03-06T19:47:25.997Z",
      "content": "<p>Since the dataset is quite big, I am planning to train the model with Spark cluster, is that okay?</p>",
      "rawMarkdown": "Since the dataset is quite big, I am planning to train the model with Spark cluster, is that okay?",
      "votes": 1,
      "replies": [
        {
          "id": 292404,
          "postDate": "2018-03-07T22:29:27.763Z",
          "content": "<p>Can we use spark cluster on kaggle?  I was thinking to use AWS for that. </p>",
          "rawMarkdown": "Can we use spark cluster on kaggle?  I was thinking to use AWS for that. "
        },
        {
          "id": 292425,
          "postDate": "2018-03-07T23:57:51.967Z",
          "content": "<p>I don't think so. Use your own cluster. </p>",
          "rawMarkdown": "I don't think so. Use your own cluster. "
        }
      ]
    },
    {
      "id": 325002,
      "postDate": "2018-05-08T03:07:59.627Z",
      "content": "<p>Hi, <a href=\"/inversion\">@inversion</a>, you should look at <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56241\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56241</a>\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56252\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56252</a>\nThe competition is terrible because of some people cheating in the last minute!</p>",
      "rawMarkdown": "Hi, @inversion, you should look at https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56241\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56252\nThe competition is terrible because of some people cheating in the last minute!",
      "votes": 2
    },
    {
      "id": 292867,
      "postDate": "2018-03-08T19:58:35.380Z",
      "content": "<p>Hi, could we get more decimals on leaderboard please? </p>",
      "rawMarkdown": "Hi, could we get more decimals on leaderboard please? ",
      "votes": 2
    },
    {
      "id": 291801,
      "postDate": "2018-03-06T21:48:51.540Z",
      "content": "<p>Hey Inversion, good to see you working for Kaggle!! This will be a nice competition.</p>",
      "rawMarkdown": "Hey Inversion, good to see you working for Kaggle!! This will be a nice competition."
    },
    {
      "id": 324193,
      "postDate": "2018-05-07T10:39:46.720Z",
      "content": "<p>why am I not able to download the data? it said downloading is currently disabled until the competition has completed.</p>",
      "rawMarkdown": "why am I not able to download the data? it said downloading is currently disabled until the competition has completed.",
      "replies": [
        {
          "id": 326470,
          "postDate": "2018-05-09T19:29:55.373Z",
          "content": "<p>Yeah, it seems they do not allow downloading data and new entry this late in competition. Once the competition is over, you will be able to download the data. However, it would mean that you would not be able to participate</p>",
          "rawMarkdown": "Yeah, it seems they do not allow downloading data and new entry this late in competition. Once the competition is over, you will be able to download the data. However, it would mean that you would not be able to participate"
        }
      ]
    },
    {
      "id": 323005,
      "postDate": "2018-05-04T06:00:39.820Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> </p>\n\n<p>I have a basic question. </p>\n\n<p>Can I submit only the submission test prediction file for Talking Data competition? </p>\n\n<p>Upload my predictions as private data set and then create a kernel to process that submission CSV? Is this acceptable under competition rules? ( I am taking this approach because of the size of the data set)</p>\n\n<p>Regards\nShanth</p>",
      "rawMarkdown": "@inversion \n\nI have a basic question. \n\nCan I submit only the submission test prediction file for Talking Data competition? \n\nUpload my predictions as private data set and then create a kernel to process that submission CSV? Is this acceptable under competition rules? ( I am taking this approach because of the size of the data set)\n\nRegards\nShanth",
      "replies": [
        {
          "id": 323310,
          "postDate": "2018-05-04T20:04:46.417Z",
          "content": "<p>As long as you're not sharing outside of a team, that's fine.</p>",
          "rawMarkdown": "As long as you're not sharing outside of a team, that's fine."
        },
        {
          "id": 323326,
          "postDate": "2018-05-04T20:50:40.923Z",
          "content": "<p>@ inversion \nThanks. I have not shared these files with anyone via any public kernels/ or outside of Kaggle. So as long as I keep my uploaded dataset files private, I should be good right? </p>\n\n<p>Regards\nPrasanth</p>",
          "rawMarkdown": "@ inversion \nThanks. I have not shared these files with anyone via any public kernels/ or outside of Kaggle. So as long as I keep my uploaded dataset files private, I should be good right? \n\nRegards\nPrasanth",
          "votes": 2
        },
        {
          "id": 323358,
          "postDate": "2018-05-04T23:10:37.300Z",
          "content": "<p>Correct.</p>",
          "rawMarkdown": "Correct."
        },
        {
          "id": 323425,
          "postDate": "2018-05-05T05:52:05.087Z",
          "content": "<p>@ Inversion. Thanks so much :)</p>",
          "rawMarkdown": "@ Inversion. Thanks so much :)"
        },
        {
          "id": 324464,
          "postDate": "2018-05-07T18:11:20.983Z",
          "content": "<p>Hi,  @Inversion, \nCan you take a look at this post: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182#324458\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182#324458</a></p>\n\n<p>Do you think there is something should be discussed and hopefully some action can be taken on Kaggle's side? The last-minute high score kernel sharing is quite discouraging to many. </p>",
          "rawMarkdown": "Hi,  @Inversion, \nCan you take a look at this post: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182#324458\n \n Do you think there is something should be discussed and hopefully some action can be taken on Kaggle's side? The last-minute high score kernel sharing is quite discouraging to many. \n\n",
          "votes": 2
        },
        {
          "id": 324951,
          "postDate": "2018-05-08T02:10:11.587Z",
          "content": "<p>hello,<a href=\"/inversion\">@inversion</a>. at the last moment of this competition,I try to get to 0.9803. I was so happy ，When I get the score. but When I get up from bed this morning，I was so upset to see others submit the same 09811 file.... sad</p>",
          "rawMarkdown": "hello,@inversion. at the last moment of this competition,I try to get to 0.9803. I was so happy ，When I get the score. but When I get up from bed this morning，I was so upset to see others submit the same 09811 file.... sad"
        }
      ]
    },
    {
      "id": 321803,
      "postDate": "2018-05-01T23:37:59.447Z",
      "content": "<p>I am unable to download data set. Do you think I am doing something wrong? It prompts that data is locked till competition is over</p>",
      "rawMarkdown": "I am unable to download data set. Do you think I am doing something wrong? It prompts that data is locked till competition is over",
      "replies": [
        {
          "id": 323309,
          "postDate": "2018-05-04T20:03:47.007Z",
          "content": "<p>The deadline to enter the contest was April 30. We don't allow new entrants the last 7 days before contest close.</p>",
          "rawMarkdown": "The deadline to enter the contest was April 30. We don't allow new entrants the last 7 days before contest close."
        },
        {
          "id": 323361,
          "postDate": "2018-05-04T23:39:16.827Z",
          "content": "<p>Thanks, I am still learning more about Kaggle and competitions </p>",
          "rawMarkdown": "Thanks, I am still learning more about Kaggle and competitions "
        }
      ]
    },
    {
      "id": 320035,
      "postDate": "2018-04-27T10:31:55.683Z",
      "content": "<p>I am working on a Rmd notebook  and  when I try to \"Commit &amp; Run\"  I am always getting the message - \"The kernel was killed for running longer than 3600 seconds.\" .Is it because of my profile restrictions ?  Can someone help ?</p>\n\n<p>UPDATE - the kernel ran succesffully , and I am able to submit to competition. There was errors \nfor some lines of code</p>",
      "rawMarkdown": "I am working on a Rmd notebook  and  when I try to \"Commit &amp; Run\"  I am always getting the message - \"The kernel was killed for running longer than 3600 seconds.\" .Is it because of my profile restrictions ?  Can someone help ?\n\n\nUPDATE - the kernel ran succesffully , and I am able to submit to competition. There was errors \nfor some lines of code"
    },
    {
      "id": 315086,
      "postDate": "2018-04-16T18:22:52.333Z",
      "content": "<p>Hyvää päivää Inversion !</p>\n\n<p>Just a general question on the data provided, as one approach is to \"follow the money\", or \"who benefits from the crime ?\".</p>\n\n<p>My understanding is that (IP defined) mobile users don't charge/invoice advertisers directly, as they do not have a business account nor entered into an agreement with any of them.</p>\n\n<p>This is the sole prerogative of the \"channel id of mobile ad publishers\", that is approx 170 of them based on @anokas, @yulia et al EDA kernels.\nThey sign deals with advertisers then push their ads on their relevant networks and invoiced them afterwards, based on agreed terms (CPM, CPC, CPA, CTR, etc.).\nSuch channels id include Google Adwords, Facebook Ads and all the major/minor players in China.</p>\n\n<p>Some may then enter into some shallow partnerships with \"mobile users\" to generate fake clicks on their networks and boost revenues, eventually paying them back via a CPC/CPA basis.</p>\n\n<p>Is that a correct understanding of the cash flow ?</p>",
      "rawMarkdown": "Hyvää päivää Inversion !\n\nJust a general question on the data provided, as one approach is to \"follow the money\", or \"who benefits from the crime ?\".\n\nMy understanding is that (IP defined) mobile users don't charge/invoice advertisers directly, as they do not have a business account nor entered into an agreement with any of them.\n\nThis is the sole prerogative of the \"channel id of mobile ad publishers\", that is approx 170 of them based on @anokas, @yulia et al EDA kernels.\nThey sign deals with advertisers then push their ads on their relevant networks and invoiced them afterwards, based on agreed terms (CPM, CPC, CPA, CTR, etc.).\nSuch channels id include Google Adwords, Facebook Ads and all the major/minor players in China.\n\nSome may then enter into some shallow partnerships with \"mobile users\" to generate fake clicks on their networks and boost revenues, eventually paying them back via a CPC/CPA basis.\n\nIs that a correct understanding of the cash flow ?",
      "replies": [
        {
          "id": 315968,
          "postDate": "2018-04-17T22:40:39.400Z",
          "content": "<p>Yes, for the most part. Although there are a few details that may or may not matter to your question. </p>\n\n<p>Generally . . .</p>\n\n<p>If I buy ads from Facebook, I pay Facebook every time there's a click (or impression). There's no incentive for click fraud.</p>\n\n<p>But what if I want to put ads on 100s of web pages or in 100s of apps? I go through an intermediary, who places my ad on sites/apps that have signed up with the intermediary. (The Google equivalent: I buy ads from Adwords. A web publisher places ads in the site using Adsense.)</p>\n\n<p>Every time someone clicks on the ad, the site/app owner (and the intermediary!) make money, at the ad buyer's expense. The more clicks, the more money. It's as easy as setting up a website, placing ads, and clicking away.</p>\n\n<p>Except, when the ad buyer figures out that the last, say, 100,000 clicks from Indonesia were almost certainly fraudulent, a refund will be requested. Now the intermediary's profit for those clicks disappears. So, it's in the intermediary's best interest to reduce click fraud.</p>\n\n<p>I think we're saying the same thing. I'm just putting it in the words I'd use to explain it.</p>",
          "rawMarkdown": "Yes, for the most part. Although there are a few details that may or may not matter to your question. \n\nGenerally . . .\n\nIf I buy ads from Facebook, I pay Facebook every time there's a click (or impression). There's no incentive for click fraud.\n\nBut what if I want to put ads on 100s of web pages or in 100s of apps? I go through an intermediary, who places my ad on sites/apps that have signed up with the intermediary. (The Google equivalent: I buy ads from Adwords. A web publisher places ads in the site using Adsense.)\n\nEvery time someone clicks on the ad, the site/app owner (and the intermediary!) make money, at the ad buyer's expense. The more clicks, the more money. It's as easy as setting up a website, placing ads, and clicking away.\n\nExcept, when the ad buyer figures out that the last, say, 100,000 clicks from Indonesia were almost certainly fraudulent, a refund will be requested. Now the intermediary's profit for those clicks disappears. So, it's in the intermediary's best interest to reduce click fraud.\n\nI think we're saying the same thing. I'm just putting it in the words I'd use to explain it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 309972,
      "postDate": "2018-04-06T10:33:03.407Z",
      "content": "<p>I have made few submissions, trying to improve my results in every next submission. Looking forward to learn a lot from community and discussion forums.</p>",
      "rawMarkdown": "I have made few submissions, trying to improve my results in every next submission. Looking forward to learn a lot from community and discussion forums."
    },
    {
      "id": 309722,
      "postDate": "2018-04-05T21:20:42.110Z",
      "content": "<p>How long does it take for a submission to go thru? it seems mine has hung up during kaggle 500 http error... @inversion any chance to look at my account?</p>",
      "rawMarkdown": "How long does it take for a submission to go thru? it seems mine has hung up during kaggle 500 http error... @inversion any chance to look at my account?"
    },
    {
      "id": 309657,
      "postDate": "2018-04-05T18:53:59.030Z",
      "content": "<p>I wish kaggle provided more than 60 min in the kernel. At least for the data sets that are this big! :( could train only on half the data due to computation limitations. Any solution anyone?</p>",
      "rawMarkdown": "I wish kaggle provided more than 60 min in the kernel. At least for the data sets that are this big! :( could train only on half the data due to computation limitations. Any solution anyone?",
      "replies": [
        {
          "id": 309823,
          "postDate": "2018-04-06T03:02:01.063Z",
          "content": "<p>I think the 60 min limitation now is only for idle time. Kernels can run for 6 hours. The memory limit of 16MB is a much bigger issue for this competition.</p>",
          "rawMarkdown": "I think the 60 min limitation now is only for idle time. Kernels can run for 6 hours. The memory limit of 16MB is a much bigger issue for this competition.",
          "votes": 2
        }
      ]
    },
    {
      "id": 308841,
      "postDate": "2018-04-04T07:23:12.033Z",
      "content": "<p>@inversion </p>\n\n<p>Will the final evaluation be done on the current test size of 18.7 Million records or is this 18% of the final data set. I am a little confused based on these instructions on the leaderboard. </p>\n\n<p>\"This leaderboard is calculated with approximately 18% of the test data.\nThe final results will be based on the other 82%, so the final standings may be different\"</p>\n\n<p>I ask since some methods I was thinking of testing out will simply not work if the final test is as big as 90 Million. Can you please confirm?</p>\n\n<p>Regards\nShanth</p>",
      "rawMarkdown": "@inversion \n\nWill the final evaluation be done on the current test size of 18.7 Million records or is this 18% of the final data set. I am a little confused based on these instructions on the leaderboard. \n\n\"This leaderboard is calculated with approximately 18% of the test data.\nThe final results will be based on the other 82%, so the final standings may be different\"\n\nI ask since some methods I was thinking of testing out will simply not work if the final test is as big as 90 Million. Can you please confirm?\n\nRegards\nShanth",
      "replies": [
        {
          "id": 309047,
          "postDate": "2018-04-04T15:00:03.083Z",
          "content": "<p>The current test set is broken into 2 parts. The first (18% of 18.7 million rows) is used to calculate the Public leaderboard. The second part (82%) is used for the final score. You make predictions for the entire set of 18.7 million, but we only show you the score for the 18% until the end of the competition. </p>",
          "rawMarkdown": "The current test set is broken into 2 parts. The first (18% of 18.7 million rows) is used to calculate the Public leaderboard. The second part (82%) is used for the final score. You make predictions for the entire set of 18.7 million, but we only show you the score for the 18% until the end of the competition. "
        },
        {
          "id": 309061,
          "postDate": "2018-04-04T15:20:41.140Z",
          "content": "<p>@ inversion </p>\n\n<p>Thanks for the prompt reply. Its clear to me now. </p>",
          "rawMarkdown": "@ inversion \n\nThanks for the prompt reply. Its clear to me now. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 297142,
      "postDate": "2018-03-16T11:37:27.840Z",
      "content": "<p>Hi Look like , I am seeing petabyte of data while extracting training data. Can you please have alook on this issue or let me know where I am going wrong??</p>",
      "rawMarkdown": "Hi Look like , I am seeing petabyte of data while extracting training data. Can you please have alook on this issue or let me know where I am going wrong??",
      "replies": [
        {
          "id": 297454,
          "postDate": "2018-03-17T04:26:39.847Z",
          "content": "<p>You'll get that error if you're using the windows default zip application. Try using an app like 7-zip to unzip the file.</p>",
          "rawMarkdown": "You'll get that error if you're using the windows default zip application. Try using an app like 7-zip to unzip the file.",
          "votes": 2
        }
      ]
    },
    {
      "id": 296109,
      "postDate": "2018-03-14T18:19:54.560Z",
      "content": "<p>Hi, is there any information about the price of user's phone or decoded phone type?</p>",
      "rawMarkdown": "Hi, is there any information about the price of user's phone or decoded phone type?"
    },
    {
      "id": 295175,
      "postDate": "2018-03-13T07:31:10.117Z",
      "content": "<p>Hi. I'm new to kagglle. How do I start a discussion? I have a question regarding the 'realness' of IP addresses. </p>\n\n<p>EDIT: Found it, please ignore this. </p>",
      "rawMarkdown": "Hi. I'm new to kagglle. How do I start a discussion? I have a question regarding the 'realness' of IP addresses. \n\nEDIT: Found it, please ignore this. "
    },
    {
      "id": 293140,
      "postDate": "2018-03-09T08:45:34.793Z",
      "content": "<p>Hello All,\nThe size of train data is 1GB (zip form). I want to understand will normal laptop process this much of data in R? \nI am using R and bit new on this. I generally worked on small dataset and my system gets hang while processing this big dataset. Can you guys guide me on this.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hello All,\nThe size of train data is 1GB (zip form). I want to understand will normal laptop process this much of data in R? \nI am using R and bit new on this. I generally worked on small dataset and my system gets hang while processing this big dataset. Can you guys guide me on this.\n\nThanks!",
      "replies": [
        {
          "id": 293517,
          "postDate": "2018-03-10T03:11:50.577Z",
          "content": "<p>it is 7GB after unzip, your laptop can process it if its memory is big enough</p>",
          "rawMarkdown": "it is 7GB after unzip, your laptop can process it if its memory is big enough"
        },
        {
          "id": 294606,
          "postDate": "2018-03-12T08:13:40.317Z",
          "content": "<p>I hope you have some good memory on it,\nI underestimated the 7.5GB file and mon Computer froze when i tried to upload it on R (with 16GB RAM), i'll retry it after a little cleaning</p>",
          "rawMarkdown": "I hope you have some good memory on it,\nI underestimated the 7.5GB file and mon Computer froze when i tried to upload it on R (with 16GB RAM), i'll retry it after a little cleaning"
        },
        {
          "id": 295627,
          "postDate": "2018-03-13T23:54:26.963Z",
          "content": "<p>you can downcast most of the variables, like in this kernel and load as much as data you can\n<a href=\"https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-val-auc-0-977\">https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-val-auc-0-977</a></p>",
          "rawMarkdown": "you can downcast most of the variables, like in this kernel and load as much as data you can\nhttps://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-val-auc-0-977",
          "votes": 1
        },
        {
          "id": 296919,
          "postDate": "2018-03-16T00:56:15.140Z",
          "content": "<p>You can have a try on the <code>ff</code> package, which uses a pointer to a flat binary file stored in the disk.\n<a href=\"https://rpubs.com/msundar/large_data_analysis\">https://rpubs.com/msundar/large_data_analysis</a></p>",
          "rawMarkdown": "You can have a try on the `ff` package, which uses a pointer to a flat binary file stored in the disk.\nhttps://rpubs.com/msundar/large_data_analysis"
        }
      ]
    },
    {
      "id": 292399,
      "postDate": "2018-03-07T22:23:36.093Z",
      "content": "<p>Is the training data a sample itself? Or is it all clicks for a period of time (over a few days)? Also, wouldn't it make more sense to build the training sample by a random selection across IP address, which might match \"individuals\" better. We're trying to predict the behavior of individuals, which means we loose a lot of inter-click correlation by sampling across clicks rather than across individuals.</p>",
      "rawMarkdown": "Is the training data a sample itself? Or is it all clicks for a period of time (over a few days)? Also, wouldn't it make more sense to build the training sample by a random selection across IP address, which might match \"individuals\" better. We're trying to predict the behavior of individuals, which means we loose a lot of inter-click correlation by sampling across clicks rather than across individuals.",
      "replies": [
        {
          "id": 292492,
          "postDate": "2018-03-08T03:18:37.350Z",
          "content": "<p>The training data is built by extract all clicks, from a random selection across IP address of interest, for a period of time over a few days. \nI didn't catch \"which means we loose a lot of inter-click correlation by sampling across clicks rather than across individuals\", maybe you explain a little more, so that I can provide whatever you need?</p>",
          "rawMarkdown": "The training data is built by extract all clicks, from a random selection across IP address of interest, for a period of time over a few days. \nI didn't catch \"which means we loose a lot of inter-click correlation by sampling across clicks rather than across individuals\", maybe you explain a little more, so that I can provide whatever you need?"
        }
      ]
    },
    {
      "id": 292307,
      "postDate": "2018-03-07T19:36:43.403Z",
      "content": "<p>this is very exciting!</p>",
      "rawMarkdown": "this is very exciting!"
    },
    {
      "id": 291870,
      "postDate": "2018-03-07T02:29:24.793Z",
      "content": "<p>I'm not sure is this the correct place to ask but according to this kernel :\n<a href=\"https://www.kaggle.com/joaopmpeinado/xgboost-starter-lb-0-944\">https://www.kaggle.com/joaopmpeinado/xgboost-starter-lb-0-944</a></p>\n\n<p>The kernel time limit seems to be 4 hours now, is that true?</p>",
      "rawMarkdown": "I'm not sure is this the correct place to ask but according to this kernel :\nhttps://www.kaggle.com/joaopmpeinado/xgboost-starter-lb-0-944\n\nThe kernel time limit seems to be 4 hours now, is that true?",
      "replies": [
        {
          "id": 292759,
          "postDate": "2018-03-08T15:24:20.643Z",
          "content": "<p>I doubt that. May be the kernel time shown is wrong!?</p>",
          "rawMarkdown": "I doubt that. May be the kernel time shown is wrong!?"
        }
      ]
    },
    {
      "id": 291344,
      "postDate": "2018-03-06T01:58:41.440Z",
      "content": "<p>;)</p>",
      "rawMarkdown": ";)"
    },
    {
      "id": 297567,
      "postDate": "2018-03-17T11:50:45.373Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 299949,
      "author_name": "harisankarh",
      "author_url": "",
      "post_date": "2018-03-21T06:27:43.990000",
      "content": "<p>There seem to be multiple questions reg IP addresses([1], [2]). @inversion would be great if you could give your inputs here. The main questions seem to be regarding the nature of the function used to anonymize the IP addresses:</p>\n\n<p>Questions are:</p>\n\n<p>(1) whether the anonymization preserves IP range similarity. i.e., if IP1 and IP2 belonged to the same class C subnet originally, will they have similar values after anonymization. (for e.g., a hash-based anonymization will not preserve this IP range similarity)</p>\n\n<p>(2) Is the same anonymization scheme used for both training and test set ? If the same IP address IP1 appears in both training and test sets before anonymization, will it have the same value in both training and test sets after anonymization ?</p>\n\n<p>[1] <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal</a></p>\n\n<p>[2] <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374</a></p>",
      "votes": 9,
      "replies": [
        {
          "id": 303477,
          "author_name": "Digvijay Rana",
          "author_url": "",
          "post_date": "2018-03-26T09:04:04.947000",
          "content": "<p>Same Question</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 291895,
      "author_name": "KALE",
      "author_url": "",
      "post_date": "2018-03-07T03:30:38.857000",
      "content": "<p>Can we get more decimal places please?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 292249,
          "author_name": "Toby Cheese",
          "author_url": "",
          "post_date": "2018-03-07T17:49:17.087000",
          "content": "<p>I second this. The leaderboard is somewhat pointless right now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 292317,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-03-07T20:02:13.543000",
          "content": "<p>Upvote for more decimals</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 292403,
          "author_name": "Harvey Pan",
          "author_url": "",
          "post_date": "2018-03-07T22:28:01.313000",
          "content": "<p>Cannot agree more.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 293297,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-03-09T15:45:24.043000",
          "content": "<p>LB now has 4 sig figs.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 324997,
      "author_name": "Shanth",
      "author_url": "",
      "post_date": "2018-05-08T03:01:28.063000",
      "content": "<p>@ inversion </p>\n\n<p>I guess the Leader board got hammered. \nLet's take a look at this <a href=\"https://www.kaggle.com/umeshsati54/merge-and-avg/data\">BLENDING OF PUBLIC CSV FILES </a></p>\n\n<p>Someone has just downloaded the outputs of several files and used the same files again. I believe these same set of files must have been used by a large number of people ( Some of the best people in the competition are nowhere to be seen )</p>\n\n<p>Is that not PLAGIARISM?????? I understand the importance of Kaggle staying open sourced but at the same time KAGGLE MUST PENALIZE plagiarism as well. </p>\n\n<p>I was so sure that this was outside the rules that I did not use the file. I didn't use that file because I thought I would get disqualified. But what I got in return was getting pushed down from top 7% to 22 % in the Leader board. </p>\n\n<p>Hoping Kaggle can do something about the final results.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 325302,
          "author_name": "Izmaylov Konstantin",
          "author_url": "",
          "post_date": "2018-05-08T09:36:49.770000",
          "content": "<p>totally agree with you! think that the very first kernel should be published and all others with 100% correlation should be threated as dublicates. So you can improve your solution based on features of this kernel and NOT to copy it blindly.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 291255,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2018-03-05T21:25:45.887000",
      "content": "<p>Thanks, @inversion! </p>\n\n<p>I wonder if it is possible for Kaggle to provide larger memory for kernels. I found it is hard to load all the data while the sampled training set is to small to fit model.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 291267,
          "author_name": "Sarthak Agarwal",
          "author_url": "",
          "post_date": "2018-03-05T21:50:59.443000",
          "content": "<p>Same question!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 291272,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-03-05T21:55:32",
          "content": "<p>Found it is ok to load 10m row but the total is ~ 185m rows  :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 291382,
          "author_name": "Muhammad Alfiansyah",
          "author_url": "",
          "post_date": "2018-03-06T03:08:51.997000",
          "content": "<p>My plan is to use data stacking and kernel chaining to beat this competition!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 291497,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-03-06T08:35:25",
          "content": "<p>@Muhammad Alfiansyah, sounds like a great plan. I am working on fit_generator from Keras so I can load/fit the training set batch by batch.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 291542,
          "author_name": "Muhammad Alfiansyah",
          "author_url": "",
          "post_date": "2018-03-06T10:22:07.473000",
          "content": "<p>@Shujian.M.Coder, looking forward to learn from and working with you!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 291671,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-03-06T16:15:30.093000",
          "content": "<p>Unfortunately, we aren't able to increase memory for Kernels in the short term.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 295862,
          "author_name": "Marissa Utterberg",
          "author_url": "",
          "post_date": "2018-03-14T10:28:56.673000",
          "content": "<p>I see that some people use AWS/GCP for additional compute capacity. Perhaps you could link to that in your Kernel? These services often provide starter credits and offer highly capable analysis platforms if you have the time to tinker with them. My guess is that Kaggle has a significantly lower budget for this than corporations like Amazon and Google, so unless we want to make Kaggle into a pay site we will have to work within the limits of digital sustainability.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 295896,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-03-14T11:37:13.877000",
          "content": "<p>Thanks, Marissa. I have almost burnt all the free credits from AWS and GCP. That's why I hope I can just run some experiments in Kaggle kernels instead pay by my own.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 326113,
      "author_name": "Shanth",
      "author_url": "",
      "post_date": "2018-05-09T09:57:40.280000",
      "content": "<p>@ inversion </p>\n\n<p>Is the Kaggle team still looking into the whole issue of shared CSV's in the Talking data competition? Are you having specific discussions internally on re-evaluating Talking Data based on the duplicate CSVs?</p>\n\n<p>OR is the competition considered to be CLOSED and no further scrutiny is to be done by Kaggle?</p>\n\n<p>Regards\nShanth</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 302624,
      "author_name": "Joe Kong",
      "author_url": "",
      "post_date": "2018-03-24T13:18:30.567000",
      "content": "<p>Why doesn't kaggle introduce joint ranking for the same scores? Doesn't seem fair that so many people can be eranked lower or higher for same score. It's the difference between winning a medal and not having one.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 317665,
          "author_name": "jiangji",
          "author_url": "",
          "post_date": "2018-04-22T07:20:17.207000",
          "content": "<p>the one achieved the score earlier has a higher rank, it is fair. and it is just based on public set anyway</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 291768,
      "author_name": "Jie Zhang",
      "author_url": "",
      "post_date": "2018-03-06T19:47:25.997000",
      "content": "<p>Since the dataset is quite big, I am planning to train the model with Spark cluster, is that okay?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 292404,
          "author_name": "Harvey Pan",
          "author_url": "",
          "post_date": "2018-03-07T22:29:27.763000",
          "content": "<p>Can we use spark cluster on kaggle?  I was thinking to use AWS for that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 292425,
          "author_name": "Jie Zhang",
          "author_url": "",
          "post_date": "2018-03-07T23:57:51.967000",
          "content": "<p>I don't think so. Use your own cluster. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325002,
      "author_name": "Shane",
      "author_url": "",
      "post_date": "2018-05-08T03:07:59.627000",
      "content": "<p>Hi, <a href=\"/inversion\">@inversion</a>, you should look at <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56241\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56241</a>\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56252\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56252</a>\nThe competition is terrible because of some people cheating in the last minute!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 292867,
      "author_name": "Snorlax",
      "author_url": "",
      "post_date": "2018-03-08T19:58:35.380000",
      "content": "<p>Hi, could we get more decimals on leaderboard please? </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 291801,
      "author_name": "Dimitris Leventis",
      "author_url": "",
      "post_date": "2018-03-06T21:48:51.540000",
      "content": "<p>Hey Inversion, good to see you working for Kaggle!! This will be a nice competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 324193,
      "author_name": "gypsysunny",
      "author_url": "",
      "post_date": "2018-05-07T10:39:46.720000",
      "content": "<p>why am I not able to download the data? it said downloading is currently disabled until the competition has completed.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 326470,
          "author_name": "Neha Bhushan",
          "author_url": "",
          "post_date": "2018-05-09T19:29:55.373000",
          "content": "<p>Yeah, it seems they do not allow downloading data and new entry this late in competition. Once the competition is over, you will be able to download the data. However, it would mean that you would not be able to participate</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 323005,
      "author_name": "Shanth",
      "author_url": "",
      "post_date": "2018-05-04T06:00:39.820000",
      "content": "<p><a href=\"/inversion\">@inversion</a> </p>\n\n<p>I have a basic question. </p>\n\n<p>Can I submit only the submission test prediction file for Talking Data competition? </p>\n\n<p>Upload my predictions as private data set and then create a kernel to process that submission CSV? Is this acceptable under competition rules? ( I am taking this approach because of the size of the data set)</p>\n\n<p>Regards\nShanth</p>",
      "votes": 0,
      "replies": [
        {
          "id": 323310,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-05-04T20:04:46.417000",
          "content": "<p>As long as you're not sharing outside of a team, that's fine.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323326,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-05-04T20:50:40.923000",
          "content": "<p>@ inversion \nThanks. I have not shared these files with anyone via any public kernels/ or outside of Kaggle. So as long as I keep my uploaded dataset files private, I should be good right? </p>\n\n<p>Regards\nPrasanth</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 323358,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-05-04T23:10:37.300000",
          "content": "<p>Correct.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323425,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-05-05T05:52:05.087000",
          "content": "<p>@ Inversion. Thanks so much :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 324464,
          "author_name": "Rui Li",
          "author_url": "",
          "post_date": "2018-05-07T18:11:20.983000",
          "content": "<p>Hi,  @Inversion, \nCan you take a look at this post: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182#324458\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56182#324458</a></p>\n\n<p>Do you think there is something should be discussed and hopefully some action can be taken on Kaggle's side? The last-minute high score kernel sharing is quite discouraging to many. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 324951,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-05-08T02:10:11.587000",
          "content": "<p>hello,<a href=\"/inversion\">@inversion</a>. at the last moment of this competition,I try to get to 0.9803. I was so happy ，When I get the score. but When I get up from bed this morning，I was so upset to see others submit the same 09811 file.... sad</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 321803,
      "author_name": "Neha Bhushan",
      "author_url": "",
      "post_date": "2018-05-01T23:37:59.447000",
      "content": "<p>I am unable to download data set. Do you think I am doing something wrong? It prompts that data is locked till competition is over</p>",
      "votes": 0,
      "replies": [
        {
          "id": 323309,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-05-04T20:03:47.007000",
          "content": "<p>The deadline to enter the contest was April 30. We don't allow new entrants the last 7 days before contest close.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323361,
          "author_name": "Neha Bhushan",
          "author_url": "",
          "post_date": "2018-05-04T23:39:16.827000",
          "content": "<p>Thanks, I am still learning more about Kaggle and competitions </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 320035,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-27T10:31:55.683000",
      "content": "<p>I am working on a Rmd notebook  and  when I try to \"Commit &amp; Run\"  I am always getting the message - \"The kernel was killed for running longer than 3600 seconds.\" .Is it because of my profile restrictions ?  Can someone help ?</p>\n\n<p>UPDATE - the kernel ran succesffully , and I am able to submit to competition. There was errors \nfor some lines of code</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 315086,
      "author_name": "Eric Perbos-Brinck",
      "author_url": "",
      "post_date": "2018-04-16T18:22:52.333000",
      "content": "<p>Hyvää päivää Inversion !</p>\n\n<p>Just a general question on the data provided, as one approach is to \"follow the money\", or \"who benefits from the crime ?\".</p>\n\n<p>My understanding is that (IP defined) mobile users don't charge/invoice advertisers directly, as they do not have a business account nor entered into an agreement with any of them.</p>\n\n<p>This is the sole prerogative of the \"channel id of mobile ad publishers\", that is approx 170 of them based on @anokas, @yulia et al EDA kernels.\nThey sign deals with advertisers then push their ads on their relevant networks and invoiced them afterwards, based on agreed terms (CPM, CPC, CPA, CTR, etc.).\nSuch channels id include Google Adwords, Facebook Ads and all the major/minor players in China.</p>\n\n<p>Some may then enter into some shallow partnerships with \"mobile users\" to generate fake clicks on their networks and boost revenues, eventually paying them back via a CPC/CPA basis.</p>\n\n<p>Is that a correct understanding of the cash flow ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 315968,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-04-17T22:40:39.400000",
          "content": "<p>Yes, for the most part. Although there are a few details that may or may not matter to your question. </p>\n\n<p>Generally . . .</p>\n\n<p>If I buy ads from Facebook, I pay Facebook every time there's a click (or impression). There's no incentive for click fraud.</p>\n\n<p>But what if I want to put ads on 100s of web pages or in 100s of apps? I go through an intermediary, who places my ad on sites/apps that have signed up with the intermediary. (The Google equivalent: I buy ads from Adwords. A web publisher places ads in the site using Adsense.)</p>\n\n<p>Every time someone clicks on the ad, the site/app owner (and the intermediary!) make money, at the ad buyer's expense. The more clicks, the more money. It's as easy as setting up a website, placing ads, and clicking away.</p>\n\n<p>Except, when the ad buyer figures out that the last, say, 100,000 clicks from Indonesia were almost certainly fraudulent, a refund will be requested. Now the intermediary's profit for those clicks disappears. So, it's in the intermediary's best interest to reduce click fraud.</p>\n\n<p>I think we're saying the same thing. I'm just putting it in the words I'd use to explain it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 309972,
      "author_name": "parul",
      "author_url": "",
      "post_date": "2018-04-06T10:33:03.407000",
      "content": "<p>I have made few submissions, trying to improve my results in every next submission. Looking forward to learn a lot from community and discussion forums.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 309722,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2018-04-05T21:20:42.110000",
      "content": "<p>How long does it take for a submission to go thru? it seems mine has hung up during kaggle 500 http error... @inversion any chance to look at my account?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 309657,
      "author_name": "Future",
      "author_url": "",
      "post_date": "2018-04-05T18:53:59.030000",
      "content": "<p>I wish kaggle provided more than 60 min in the kernel. At least for the data sets that are this big! :( could train only on half the data due to computation limitations. Any solution anyone?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 309823,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-04-06T03:02:01.063000",
          "content": "<p>I think the 60 min limitation now is only for idle time. Kernels can run for 6 hours. The memory limit of 16MB is a much bigger issue for this competition.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 308841,
      "author_name": "Shanth",
      "author_url": "",
      "post_date": "2018-04-04T07:23:12.033000",
      "content": "<p>@inversion </p>\n\n<p>Will the final evaluation be done on the current test size of 18.7 Million records or is this 18% of the final data set. I am a little confused based on these instructions on the leaderboard. </p>\n\n<p>\"This leaderboard is calculated with approximately 18% of the test data.\nThe final results will be based on the other 82%, so the final standings may be different\"</p>\n\n<p>I ask since some methods I was thinking of testing out will simply not work if the final test is as big as 90 Million. Can you please confirm?</p>\n\n<p>Regards\nShanth</p>",
      "votes": 0,
      "replies": [
        {
          "id": 309047,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-04-04T15:00:03.083000",
          "content": "<p>The current test set is broken into 2 parts. The first (18% of 18.7 million rows) is used to calculate the Public leaderboard. The second part (82%) is used for the final score. You make predictions for the entire set of 18.7 million, but we only show you the score for the 18% until the end of the competition. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309061,
          "author_name": "Shanth",
          "author_url": "",
          "post_date": "2018-04-04T15:20:41.140000",
          "content": "<p>@ inversion </p>\n\n<p>Thanks for the prompt reply. Its clear to me now. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 297142,
      "author_name": "Suyash Mishra",
      "author_url": "",
      "post_date": "2018-03-16T11:37:27.840000",
      "content": "<p>Hi Look like , I am seeing petabyte of data while extracting training data. Can you please have alook on this issue or let me know where I am going wrong??</p>",
      "votes": 0,
      "replies": [
        {
          "id": 297454,
          "author_name": "Macsmith",
          "author_url": "",
          "post_date": "2018-03-17T04:26:39.847000",
          "content": "<p>You'll get that error if you're using the windows default zip application. Try using an app like 7-zip to unzip the file.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 296109,
      "author_name": "JiaYiZhang",
      "author_url": "",
      "post_date": "2018-03-14T18:19:54.560000",
      "content": "<p>Hi, is there any information about the price of user's phone or decoded phone type?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 295175,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-13T07:31:10.117000",
      "content": "<p>Hi. I'm new to kagglle. How do I start a discussion? I have a question regarding the 'realness' of IP addresses. </p>\n\n<p>EDIT: Found it, please ignore this. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 293140,
      "author_name": "Vaibhav",
      "author_url": "",
      "post_date": "2018-03-09T08:45:34.793000",
      "content": "<p>Hello All,\nThe size of train data is 1GB (zip form). I want to understand will normal laptop process this much of data in R? \nI am using R and bit new on this. I generally worked on small dataset and my system gets hang while processing this big dataset. Can you guys guide me on this.</p>\n\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 293517,
          "author_name": "Zootojia",
          "author_url": "",
          "post_date": "2018-03-10T03:11:50.577000",
          "content": "<p>it is 7GB after unzip, your laptop can process it if its memory is big enough</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 294606,
          "author_name": "Harmanax",
          "author_url": "",
          "post_date": "2018-03-12T08:13:40.317000",
          "content": "<p>I hope you have some good memory on it,\nI underestimated the 7.5GB file and mon Computer froze when i tried to upload it on R (with 16GB RAM), i'll retry it after a little cleaning</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 295627,
          "author_name": "Ravi Teja Gutta",
          "author_url": "",
          "post_date": "2018-03-13T23:54:26.963000",
          "content": "<p>you can downcast most of the variables, like in this kernel and load as much as data you can\n<a href=\"https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-val-auc-0-977\">https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-val-auc-0-977</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 296919,
          "author_name": "Xiyang",
          "author_url": "",
          "post_date": "2018-03-16T00:56:15.140000",
          "content": "<p>You can have a try on the <code>ff</code> package, which uses a pointer to a flat binary file stored in the disk.\n<a href=\"https://rpubs.com/msundar/large_data_analysis\">https://rpubs.com/msundar/large_data_analysis</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 292399,
      "author_name": "Jim C. March",
      "author_url": "",
      "post_date": "2018-03-07T22:23:36.093000",
      "content": "<p>Is the training data a sample itself? Or is it all clicks for a period of time (over a few days)? Also, wouldn't it make more sense to build the training sample by a random selection across IP address, which might match \"individuals\" better. We're trying to predict the behavior of individuals, which means we loose a lot of inter-click correlation by sampling across clicks rather than across individuals.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 292492,
          "author_name": "Aaron Yin",
          "author_url": "",
          "post_date": "2018-03-08T03:18:37.350000",
          "content": "<p>The training data is built by extract all clicks, from a random selection across IP address of interest, for a period of time over a few days. \nI didn't catch \"which means we loose a lot of inter-click correlation by sampling across clicks rather than across individuals\", maybe you explain a little more, so that I can provide whatever you need?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 292307,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-07T19:36:43.403000",
      "content": "<p>this is very exciting!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 291870,
      "author_name": "Muhammad Alfiansyah",
      "author_url": "",
      "post_date": "2018-03-07T02:29:24.793000",
      "content": "<p>I'm not sure is this the correct place to ask but according to this kernel :\n<a href=\"https://www.kaggle.com/joaopmpeinado/xgboost-starter-lb-0-944\">https://www.kaggle.com/joaopmpeinado/xgboost-starter-lb-0-944</a></p>\n\n<p>The kernel time limit seems to be 4 hours now, is that true?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 292759,
          "author_name": "Harvey Pan",
          "author_url": "",
          "post_date": "2018-03-08T15:24:20.643000",
          "content": "<p>I doubt that. May be the kernel time shown is wrong!?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 291344,
      "author_name": "David Z",
      "author_url": "",
      "post_date": "2018-03-06T01:58:41.440000",
      "content": "<p>;)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 297567,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-17T11:50:45.373000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "291215": "Welcome to the TalkingData AdTracking Fraud Detection Challenge! This is a challenge for those who love working on large, imbalanced, binary classification problems. (Or for those who want to build their skills in that area!)\n\nNote that the training data is ~8 Gb uncompressed. For those who want a quick preview, we've added a training sample file that contains 100k random rows. \n\nPlease feel free to ask your questions in this thread.\n\nGood luck!",
    "299949": "There seem to be multiple questions reg IP addresses([1], [2]). @inversion would be great if you could give your inputs here. The main questions seem to be regarding the nature of the function used to anonymize the IP addresses:\n\nQuestions are:\n\n(1) whether the anonymization preserves IP range similarity. i.e., if IP1 and IP2 belonged to the same class C subnet originally, will they have similar values after anonymization. (for e.g., a hash-based anonymization will not preserve this IP range similarity)\n\n(2) Is the same anonymization scheme used for both training and test set ? If the same IP address IP1 appears in both training and test sets before anonymization, will it have the same value in both training and test sets after anonymization ?\n \n[1] https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\n\n[2] https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52374",
    "291895": "Can we get more decimal places please?",
    "324997": "@ inversion \n\nI guess the Leader board got hammered. \nLet's take a look at this [BLENDING OF PUBLIC CSV FILES ][1]\n\nSomeone has just downloaded the outputs of several files and used the same files again. I believe these same set of files must have been used by a large number of people ( Some of the best people in the competition are nowhere to be seen )\n\nIs that not PLAGIARISM?????? I understand the importance of Kaggle staying open sourced but at the same time KAGGLE MUST PENALIZE plagiarism as well. \n\nI was so sure that this was outside the rules that I did not use the file. I didn't use that file because I thought I would get disqualified. But what I got in return was getting pushed down from top 7% to 22 % in the Leader board. \n\nHoping Kaggle can do something about the final results.\n\n  [1]: https://www.kaggle.com/umeshsati54/merge-and-avg/data",
    "291255": "Thanks, @inversion! \n\nI wonder if it is possible for Kaggle to provide larger memory for kernels. I found it is hard to load all the data while the sampled training set is to small to fit model.",
    "326113": "@ inversion \n\nIs the Kaggle team still looking into the whole issue of shared CSV's in the Talking data competition? Are you having specific discussions internally on re-evaluating Talking Data based on the duplicate CSVs?\n\n OR is the competition considered to be CLOSED and no further scrutiny is to be done by Kaggle?\n\nRegards\nShanth",
    "302624": "Why doesn't kaggle introduce joint ranking for the same scores? Doesn't seem fair that so many people can be eranked lower or higher for same score. It's the difference between winning a medal and not having one.",
    "291768": "Since the dataset is quite big, I am planning to train the model with Spark cluster, is that okay?",
    "325002": "Hi, @inversion, you should look at https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56241\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56252\nThe competition is terrible because of some people cheating in the last minute!",
    "292867": "Hi, could we get more decimals on leaderboard please? ",
    "291801": "Hey Inversion, good to see you working for Kaggle!! This will be a nice competition.",
    "324193": "why am I not able to download the data? it said downloading is currently disabled until the competition has completed.",
    "323005": "@inversion \n\nI have a basic question. \n\nCan I submit only the submission test prediction file for Talking Data competition? \n\nUpload my predictions as private data set and then create a kernel to process that submission CSV? Is this acceptable under competition rules? ( I am taking this approach because of the size of the data set)\n\nRegards\nShanth",
    "321803": "I am unable to download data set. Do you think I am doing something wrong? It prompts that data is locked till competition is over",
    "320035": "I am working on a Rmd notebook  and  when I try to \"Commit &amp; Run\"  I am always getting the message - \"The kernel was killed for running longer than 3600 seconds.\" .Is it because of my profile restrictions ?  Can someone help ?\n\n\nUPDATE - the kernel ran succesffully , and I am able to submit to competition. There was errors \nfor some lines of code",
    "315086": "Hyvää päivää Inversion !\n\nJust a general question on the data provided, as one approach is to \"follow the money\", or \"who benefits from the crime ?\".\n\nMy understanding is that (IP defined) mobile users don't charge/invoice advertisers directly, as they do not have a business account nor entered into an agreement with any of them.\n\nThis is the sole prerogative of the \"channel id of mobile ad publishers\", that is approx 170 of them based on @anokas, @yulia et al EDA kernels.\nThey sign deals with advertisers then push their ads on their relevant networks and invoiced them afterwards, based on agreed terms (CPM, CPC, CPA, CTR, etc.).\nSuch channels id include Google Adwords, Facebook Ads and all the major/minor players in China.\n\nSome may then enter into some shallow partnerships with \"mobile users\" to generate fake clicks on their networks and boost revenues, eventually paying them back via a CPC/CPA basis.\n\nIs that a correct understanding of the cash flow ?",
    "309972": "I have made few submissions, trying to improve my results in every next submission. Looking forward to learn a lot from community and discussion forums.",
    "309722": "How long does it take for a submission to go thru? it seems mine has hung up during kaggle 500 http error... @inversion any chance to look at my account?",
    "309657": "I wish kaggle provided more than 60 min in the kernel. At least for the data sets that are this big! :( could train only on half the data due to computation limitations. Any solution anyone?",
    "308841": "@inversion \n\nWill the final evaluation be done on the current test size of 18.7 Million records or is this 18% of the final data set. I am a little confused based on these instructions on the leaderboard. \n\n\"This leaderboard is calculated with approximately 18% of the test data.\nThe final results will be based on the other 82%, so the final standings may be different\"\n\nI ask since some methods I was thinking of testing out will simply not work if the final test is as big as 90 Million. Can you please confirm?\n\nRegards\nShanth",
    "297142": "Hi Look like , I am seeing petabyte of data while extracting training data. Can you please have alook on this issue or let me know where I am going wrong??",
    "296109": "Hi, is there any information about the price of user's phone or decoded phone type?",
    "295175": "Hi. I'm new to kagglle. How do I start a discussion? I have a question regarding the 'realness' of IP addresses. \n\nEDIT: Found it, please ignore this. ",
    "293140": "Hello All,\nThe size of train data is 1GB (zip form). I want to understand will normal laptop process this much of data in R? \nI am using R and bit new on this. I generally worked on small dataset and my system gets hang while processing this big dataset. Can you guys guide me on this.\n\nThanks!",
    "292399": "Is the training data a sample itself? Or is it all clicks for a period of time (over a few days)? Also, wouldn't it make more sense to build the training sample by a random selection across IP address, which might match \"individuals\" better. We're trying to predict the behavior of individuals, which means we loose a lot of inter-click correlation by sampling across clicks rather than across individuals.",
    "292307": "this is very exciting!",
    "291870": "I'm not sure is this the correct place to ask but according to this kernel :\nhttps://www.kaggle.com/joaopmpeinado/xgboost-starter-lb-0-944\n\nThe kernel time limit seems to be 4 hours now, is that true?",
    "291344": ";)",
    "297567": ""
  }
}