{
  "id": 51411,
  "title": "How to manage the amount of data",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51411",
  "author_name": "Asparuh Hristov",
  "post_date": "2018-03-08T15:21:01.485000",
  "votes": 69,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Hereby, I have gathered ideas on how to manage the huge amount of data without dropping (or not much) the predictive power. I will start with a very good observation from Konrad Banachewicz:</p>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347</a>\nNamely, get rid of the year and month as they are all the same</li>\n<li>I will propose to get rid of the minutes and seconds as well. One can still detect clicks coming from the same IP within an \"unreasonable\" amount of time by just comparing the date + hour. Even in the beginning simply using only the hour will be more than good enough.</li>\n<li>use dask instead of pandas (way faster!!!!). I have also just read about \"pandas on ray\", have never used it myself but looks extremely promising.</li>\n<li>in my opinion the column \"attributed time\" is totally useless - don't even read it when loading the files (the train data). What does it matter when it was attributed if you have to predict whether it is attributed or not?</li>\n<li>Count the clicks from each IP and store this information (as a person with experience in online advertisement I can guarantee you that this will be one of the most influential features). Afterwards delete IP.</li>\n<li>Actually before deleting it you can also create a column indicating if the user made more than x (say, 5) clicks within an hour.</li>\n<li>Use the correct datatypes when loading the data</li>\n<li>Use gc (garbage collector). Extremely easy to use and very helpful. Constantly delete anything that is not necessary anymore.</li>\n</ol>\n\n<p>Please, share more ideas/observations if you can think of smt.</p>",
  "messages": [
    {
      "id": 292758,
      "postDate": "2018-03-08T15:21:01.487Z",
      "content": "<p>Hereby, I have gathered ideas on how to manage the huge amount of data without dropping (or not much) the predictive power. I will start with a very good observation from Konrad Banachewicz:</p>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347</a>\nNamely, get rid of the year and month as they are all the same</li>\n<li>I will propose to get rid of the minutes and seconds as well. One can still detect clicks coming from the same IP within an \"unreasonable\" amount of time by just comparing the date + hour. Even in the beginning simply using only the hour will be more than good enough.</li>\n<li>use dask instead of pandas (way faster!!!!). I have also just read about \"pandas on ray\", have never used it myself but looks extremely promising.</li>\n<li>in my opinion the column \"attributed time\" is totally useless - don't even read it when loading the files (the train data). What does it matter when it was attributed if you have to predict whether it is attributed or not?</li>\n<li>Count the clicks from each IP and store this information (as a person with experience in online advertisement I can guarantee you that this will be one of the most influential features). Afterwards delete IP.</li>\n<li>Actually before deleting it you can also create a column indicating if the user made more than x (say, 5) clicks within an hour.</li>\n<li>Use the correct datatypes when loading the data</li>\n<li>Use gc (garbage collector). Extremely easy to use and very helpful. Constantly delete anything that is not necessary anymore.</li>\n</ol>\n\n<p>Please, share more ideas/observations if you can think of smt.</p>",
      "rawMarkdown": "Hereby, I have gathered ideas on how to manage the huge amount of data without dropping (or not much) the predictive power. I will start with a very good observation from Konrad Banachewicz:\n\n 1. https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347\nNamely, get rid of the year and month as they are all the same\n 2. I will propose to get rid of the minutes and seconds as well. One can still detect clicks coming from the same IP within an \"unreasonable\" amount of time by just comparing the date + hour. Even in the beginning simply using only the hour will be more than good enough.\n 3. use dask instead of pandas (way faster!!!!). I have also just read about \"pandas on ray\", have never used it myself but looks extremely promising.\n 4. in my opinion the column \"attributed time\" is totally useless - don't even read it when loading the files (the train data). What does it matter when it was attributed if you have to predict whether it is attributed or not?\n 5. Count the clicks from each IP and store this information (as a person with experience in online advertisement I can guarantee you that this will be one of the most influential features). Afterwards delete IP.\n 6. Actually before deleting it you can also create a column indicating if the user made more than x (say, 5) clicks within an hour.\n 7. Use the correct datatypes when loading the data\n 8. Use gc (garbage collector). Extremely easy to use and very helpful. Constantly delete anything that is not necessary anymore.\n\nPlease, share more ideas/observations if you can think of smt.",
      "votes": 69
    },
    {
      "id": 294940,
      "postDate": "2018-03-12T21:24:19.960Z",
      "content": "<p>I created a tutorial kernel with some tips on how to manage data.  It includes subsampling, loading in chunks,  creative processing and basic dask tutorial. <br>\n<a href=\"https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\">https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask</a>. <br>\nHope you guys find it useful!</p>",
      "rawMarkdown": "I created a tutorial kernel with some tips on how to manage data.  It includes subsampling, loading in chunks,  creative processing and basic dask tutorial.  \nhttps://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask.   \nHope you guys find it useful!\n\n",
      "votes": 10,
      "replies": [
        {
          "id": 298519,
          "postDate": "2018-03-19T16:56:02.907Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 295308,
      "postDate": "2018-03-13T13:15:43.847Z",
      "content": "<p>Step 2 i.e. removal of attribute_time column and keeping click_time hr and day. I have done this using python and sqlite3 (inbuilt package of python). Code of the same is here :- <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51823\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51823</a></p>\n\n<p>Feel free to use this. I think after implementing the above stage, data size can be reduced to around 5 Gb which i think is quite manageable. </p>",
      "rawMarkdown": "Step 2 i.e. removal of attribute_time column and keeping click_time hr and day. I have done this using python and sqlite3 (inbuilt package of python). Code of the same is here :- https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51823\n\nFeel free to use this. I think after implementing the above stage, data size can be reduced to around 5 Gb which i think is quite manageable. ",
      "votes": 4
    },
    {
      "id": 303115,
      "postDate": "2018-03-25T14:55:55.867Z",
      "content": "<p><strong>'the column \"attributed time\" is totally useless'</strong>\nI do not agree with this statement. How to know a priori whether the IPs in which that variable is a date one day after the last click are always reliable?\nThis is only a hypothetical example but I want to point out that it is algorithms and construction techniques that must decide whether a variable is important or not.\nObviously the intuition of the data scientist is important but I think it is good practice to verify the importance of having a variable within the model before eliminating it.</p>",
      "rawMarkdown": "**'the column \"attributed time\" is totally useless'**\nI do not agree with this statement. How to know a priori whether the IPs in which that variable is a date one day after the last click are always reliable?\nThis is only a hypothetical example but I want to point out that it is algorithms and construction techniques that must decide whether a variable is important or not.\nObviously the intuition of the data scientist is important but I think it is good practice to verify the importance of having a variable within the model before eliminating it.",
      "votes": 1
    },
    {
      "id": 296131,
      "postDate": "2018-03-14T19:01:03.620Z",
      "content": "<p>I do not have enough computing resources to participate in this competition, 16 GB Ram, 16 GB Swap(SSD), even after using data size reduction techniques discussed  in discussions I am not able to complete a full training pass on all data set. <br> I am planing to find a scheme to select a subset of train and validation data from whole train set and use that for model building through out the competition. Any suggestions?</p>",
      "rawMarkdown": "I do not have enough computing resources to participate in this competition, 16 GB Ram, 16 GB Swap(SSD), even after using data size reduction techniques discussed  in discussions I am not able to complete a full training pass on all data set. <br> I am planing to find a scheme to select a subset of train and validation data from whole train set and use that for model building through out the competition. Any suggestions?",
      "votes": 1,
      "replies": [
        {
          "id": 303443,
          "postDate": "2018-03-26T07:35:26.980Z",
          "content": "<p>You should take a look at this excellent kernel of @Yulia :</p>\n\n<p><a href=\"https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\">https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask</a>. </p>",
          "rawMarkdown": "You should take a look at this excellent kernel of @Yulia :\n\nhttps://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 293391,
      "postDate": "2018-03-09T19:17:39.630Z",
      "content": "<p>Thanks for the shout-out - glad you liked the idea  :-)</p>",
      "rawMarkdown": "Thanks for the shout-out - glad you liked the idea  :-)",
      "votes": 1
    },
    {
      "id": 293073,
      "postDate": "2018-03-09T05:46:44.870Z",
      "content": "<p>I Tried loading this via pyspark ... it does a decent job, Mllib does most of weightlifting of Random forest... </p>",
      "rawMarkdown": "I Tried loading this via pyspark ... it does a decent job, Mllib does most of weightlifting of Random forest... ",
      "votes": 1
    },
    {
      "id": 293009,
      "postDate": "2018-03-09T02:54:05.380Z",
      "content": "<p>good insights. Are u ruling out the possibility of shared IP? I mean, multiple users with same IP can click on the ad without fraudulent intent, ryt?</p>",
      "rawMarkdown": "good insights. Are u ruling out the possibility of shared IP? I mean, multiple users with same IP can click on the ad without fraudulent intent, ryt?",
      "votes": 1,
      "replies": [
        {
          "id": 293143,
          "postDate": "2018-03-09T08:55:43.403Z",
          "content": "<p>Indeed, one should think of those things, my idea was just to share general guidelines. At further stage this can be tackled in not so complicated ways - for example instead of aggregating solely on IP level , group by the (IP, app, device/os) tuple. </p>",
          "rawMarkdown": "Indeed, one should think of those things, my idea was just to share general guidelines. At further stage this can be tackled in not so complicated ways - for example instead of aggregating solely on IP level , group by the (IP, app, device/os) tuple. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 292783,
      "postDate": "2018-03-08T15:55:49.623Z",
      "content": "<p>Negative downsampling (i.e downsample the majority class). We're looking for the behavior of a small subset, not the general pop (especially with data anonymized). \n<a href=\"https://www.kaggle.com/danofer/downsampling-for-fun-speed\">https://www.kaggle.com/danofer/downsampling-for-fun-speed</a></p>\n\n<p>I've had excellent results with this approach in RW problems involving more data, albeit non anonymized problems (and with corporate resources/servers handy).</p>",
      "rawMarkdown": "Negative downsampling (i.e downsample the majority class). We're looking for the behavior of a small subset, not the general pop (especially with data anonymized). \nhttps://www.kaggle.com/danofer/downsampling-for-fun-speed\n\nI've had excellent results with this approach in RW problems involving more data, albeit non anonymized problems (and with corporate resources/servers handy).",
      "votes": 2,
      "replies": [
        {
          "id": 292788,
          "postDate": "2018-03-08T16:00:10.443Z",
          "content": "<p>Agree. I believe a good downsampling approach would be to get rid of x% of the IP's that had only a single click and not conversion (is_attributed == 0). Those samples carry pretty much the least useful information...</p>",
          "rawMarkdown": "Agree. I believe a good downsampling approach would be to get rid of x% of the IP's that had only a single click and not conversion (is_attributed == 0). Those samples carry pretty much the least useful information...",
          "votes": 4
        },
        {
          "id": 292794,
          "postDate": "2018-03-08T16:17:11.957Z",
          "content": "<p>Interesting idea. !\nI'd considered more \"directed\" downsampling (e.g. keep only IPs in the test set or is_attributed == 1 , but then there's a strong risk of bias/shift, e.g. in the other features) </p>",
          "rawMarkdown": "Interesting idea. !\nI'd considered more \"directed\" downsampling (e.g. keep only IPs in the test set or is_attributed == 1 , but then there's a strong risk of bias/shift, e.g. in the other features) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 304699,
      "postDate": "2018-03-27T22:35:56.957Z",
      "content": "<p>Are you guys training with all the data? If not, how will you choose the subsample? </p>\n\n<p>Moreover, is it worth splitting the dataset in a way that the validation is closer time-wise to the test set? Or is it better to do a stratified separation?</p>",
      "rawMarkdown": "Are you guys training with all the data? If not, how will you choose the subsample? \n\nMoreover, is it worth splitting the dataset in a way that the validation is closer time-wise to the test set? Or is it better to do a stratified separation?"
    },
    {
      "id": 303078,
      "postDate": "2018-03-25T13:03:33.377Z",
      "content": "<p>Thank you for your wonderful ideas but I have yet to manage this dataset.\nOnce I load the dataset for xgboost (DMatrix function), that volume get bigger.\nI think it is useful to save memory to use libsvm-txt format because it can be directly loaded into xgboost but it's difficult to convert.</p>",
      "rawMarkdown": "Thank you for your wonderful ideas but I have yet to manage this dataset.\nOnce I load the dataset for xgboost (DMatrix function), that volume get bigger.\nI think it is useful to save memory to use libsvm-txt format because it can be directly loaded into xgboost but it's difficult to convert."
    },
    {
      "id": 299940,
      "postDate": "2018-03-21T06:15:56.710Z",
      "content": "<p>I am just an undergraduate student at iit (ism) , dhanbad , india. How long will it take me to learn everything?</p>",
      "rawMarkdown": "I am just an undergraduate student at iit (ism) , dhanbad , india. How long will it take me to learn everything?",
      "replies": [
        {
          "id": 299946,
          "postDate": "2018-03-21T06:20:53.997Z",
          "content": "<p>You can never learn everything in data science... Just do it regualary and enjoy</p>",
          "rawMarkdown": "You can never learn everything in data science... Just do it regualary and enjoy",
          "votes": 3
        }
      ]
    },
    {
      "id": 297458,
      "postDate": "2018-03-17T04:35:20.330Z",
      "content": "<p>Here is another one, we can use out-of-core learning algos(vowpal wabbit), we can use algos that use stochastic gradient descent(logistic regression, neural nets)  </p>",
      "rawMarkdown": "Here is another one, we can use out-of-core learning algos(vowpal wabbit), we can use algos that use stochastic gradient descent(logistic regression, neural nets)  "
    },
    {
      "id": 296499,
      "postDate": "2018-03-15T10:16:51.537Z",
      "content": "<p>is there any packages/approach suggested using R?</p>",
      "rawMarkdown": "is there any packages/approach suggested using R?",
      "replies": [
        {
          "id": 296502,
          "postDate": "2018-03-15T10:18:12.800Z",
          "content": "<p>As a starting point, I would say data.table for aggregation / cleanup / processing, after that lightgbm for modeling.</p>",
          "rawMarkdown": "As a starting point, I would say data.table for aggregation / cleanup / processing, after that lightgbm for modeling.",
          "votes": 5
        }
      ]
    },
    {
      "id": 295970,
      "postDate": "2018-03-14T13:57:50.323Z",
      "content": "<p>Good Insights.. I'm new to the field of Data Science, I searched a few datasets those were many hundreds GB (e.g. 800Gb, 900GB etc). How can we work on such datasets when we didn't have that much storage space or comparatively slower Internet connections. What are cloud resources we can utilize with cheaper costs.?</p>",
      "rawMarkdown": "Good Insights.. I'm new to the field of Data Science, I searched a few datasets those were many hundreds GB (e.g. 800Gb, 900GB etc). How can we work on such datasets when we didn't have that much storage space or comparatively slower Internet connections. What are cloud resources we can utilize with cheaper costs.?\n\n\n\n"
    },
    {
      "id": 294650,
      "postDate": "2018-03-12T11:10:36.890Z",
      "content": "<p>\"Count the clicks from each IP and store this information (as a person with experience in online advertisement I can guarantee you that this will be one of the most influential features). Afterwards delete IP.\"</p>\n\n<p>Which will be more useful information:- \n1. calculating this information using both train and test.\n2. calculating separately for test and train</p>\n\n<p>If I do it like the first option, what type of biases it will create?</p>",
      "rawMarkdown": "\"Count the clicks from each IP and store this information (as a person with experience in online advertisement I can guarantee you that this will be one of the most influential features). Afterwards delete IP.\"\n\nWhich will be more useful information:- \n1. calculating this information using both train and test.\n2. calculating separately for test and train\n\nIf I do it like the first option, what type of biases it will create?",
      "replies": [
        {
          "id": 294877,
          "postDate": "2018-03-12T18:54:57.180Z",
          "content": "<p>I would say number of clicks until conversion</p>",
          "rawMarkdown": "I would say number of clicks until conversion"
        }
      ]
    },
    {
      "id": 294118,
      "postDate": "2018-03-11T10:17:27.680Z",
      "content": "<p>Thanks for this tips ! </p>\n\n<p>I tried to use dask but when if we use sklearn or XGBoost algorithms, we have to use pandas so it's usefull just for data processing !</p>",
      "rawMarkdown": "Thanks for this tips ! \n\nI tried to use dask but when if we use sklearn or XGBoost algorithms, we have to use pandas so it's usefull just for data processing !"
    },
    {
      "id": 293575,
      "postDate": "2018-03-10T06:55:14.317Z",
      "content": "<p>hello, i'm new to kaggle and datascience, i wonder how i could use this huge volume of data, when i have a poor pc and internet connection. is there a way to run my script in cloud.</p>",
      "rawMarkdown": "hello, i'm new to kaggle and datascience, i wonder how i could use this huge volume of data, when i have a poor pc and internet connection. is there a way to run my script in cloud.",
      "replies": [
        {
          "id": 293598,
          "postDate": "2018-03-10T07:47:46.907Z",
          "content": "<p>Hey Zoubair, welcome to kaggle, There are kernel that kaggle offers which you could use, it burns cpus held by kaggle &amp; wont put a load on your pc</p>",
          "rawMarkdown": "Hey Zoubair, welcome to kaggle, There are kernel that kaggle offers which you could use, it burns cpus held by kaggle &amp; wont put a load on your pc",
          "votes": 1
        },
        {
          "id": 294221,
          "postDate": "2018-03-11T15:36:14.563Z",
          "content": "<p>thank you.</p>",
          "rawMarkdown": "thank you."
        }
      ]
    },
    {
      "id": 293308,
      "postDate": "2018-03-09T16:03:51.323Z",
      "content": "<p>Thanks for sharing all this information, this could be very helpful to everyone.</p>",
      "rawMarkdown": "Thanks for sharing all this information, this could be very helpful to everyone."
    },
    {
      "id": 300558,
      "postDate": "2018-03-21T18:03:23.157Z",
      "content": "<p>Thanks for the insights</p>",
      "rawMarkdown": "Thanks for the insights"
    },
    {
      "id": 296651,
      "postDate": "2018-03-15T15:08:37.663Z",
      "content": "<p>wonderful idea, thanks!</p>",
      "rawMarkdown": "wonderful idea, thanks!"
    },
    {
      "id": 295595,
      "postDate": "2018-03-13T22:03:45.533Z",
      "content": "<p>Thanks! This is useful.</p>",
      "rawMarkdown": "Thanks! This is useful.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 294940,
      "author_name": "yulia",
      "author_url": "",
      "post_date": "2018-03-12T21:24:19.960000",
      "content": "<p>I created a tutorial kernel with some tips on how to manage data.  It includes subsampling, loading in chunks,  creative processing and basic dask tutorial. <br>\n<a href=\"https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\">https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask</a>. <br>\nHope you guys find it useful!</p>",
      "votes": 10,
      "replies": [
        {
          "id": 298519,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-03-19T16:56:02.907000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 295308,
      "author_name": "Prince Atul",
      "author_url": "",
      "post_date": "2018-03-13T13:15:43.847000",
      "content": "<p>Step 2 i.e. removal of attribute_time column and keeping click_time hr and day. I have done this using python and sqlite3 (inbuilt package of python). Code of the same is here :- <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51823\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51823</a></p>\n\n<p>Feel free to use this. I think after implementing the above stage, data size can be reduced to around 5 Gb which i think is quite manageable. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 303115,
      "author_name": "ddrbcn",
      "author_url": "",
      "post_date": "2018-03-25T14:55:55.867000",
      "content": "<p><strong>'the column \"attributed time\" is totally useless'</strong>\nI do not agree with this statement. How to know a priori whether the IPs in which that variable is a date one day after the last click are always reliable?\nThis is only a hypothetical example but I want to point out that it is algorithms and construction techniques that must decide whether a variable is important or not.\nObviously the intuition of the data scientist is important but I think it is good practice to verify the importance of having a variable within the model before eliminating it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 296131,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-03-14T19:01:03.620000",
      "content": "<p>I do not have enough computing resources to participate in this competition, 16 GB Ram, 16 GB Swap(SSD), even after using data size reduction techniques discussed  in discussions I am not able to complete a full training pass on all data set. <br> I am planing to find a scheme to select a subset of train and validation data from whole train set and use that for model building through out the competition. Any suggestions?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 303443,
          "author_name": "BenMelloul",
          "author_url": "",
          "post_date": "2018-03-26T07:35:26.980000",
          "content": "<p>You should take a look at this excellent kernel of @Yulia :</p>\n\n<p><a href=\"https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\">https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask</a>. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 293391,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2018-03-09T19:17:39.630000",
      "content": "<p>Thanks for the shout-out - glad you liked the idea  :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 293073,
      "author_name": "tarun",
      "author_url": "",
      "post_date": "2018-03-09T05:46:44.870000",
      "content": "<p>I Tried loading this via pyspark ... it does a decent job, Mllib does most of weightlifting of Random forest... </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 293009,
      "author_name": "Nagaraju Oruganti",
      "author_url": "",
      "post_date": "2018-03-09T02:54:05.380000",
      "content": "<p>good insights. Are u ruling out the possibility of shared IP? I mean, multiple users with same IP can click on the ad without fraudulent intent, ryt?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 293143,
          "author_name": "Asparuh Hristov",
          "author_url": "",
          "post_date": "2018-03-09T08:55:43.403000",
          "content": "<p>Indeed, one should think of those things, my idea was just to share general guidelines. At further stage this can be tackled in not so complicated ways - for example instead of aggregating solely on IP level , group by the (IP, app, device/os) tuple. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 292783,
      "author_name": "Dan Ofer",
      "author_url": "",
      "post_date": "2018-03-08T15:55:49.623000",
      "content": "<p>Negative downsampling (i.e downsample the majority class). We're looking for the behavior of a small subset, not the general pop (especially with data anonymized). \n<a href=\"https://www.kaggle.com/danofer/downsampling-for-fun-speed\">https://www.kaggle.com/danofer/downsampling-for-fun-speed</a></p>\n\n<p>I've had excellent results with this approach in RW problems involving more data, albeit non anonymized problems (and with corporate resources/servers handy).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 292788,
          "author_name": "Asparuh Hristov",
          "author_url": "",
          "post_date": "2018-03-08T16:00:10.443000",
          "content": "<p>Agree. I believe a good downsampling approach would be to get rid of x% of the IP's that had only a single click and not conversion (is_attributed == 0). Those samples carry pretty much the least useful information...</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 292794,
          "author_name": "Dan Ofer",
          "author_url": "",
          "post_date": "2018-03-08T16:17:11.957000",
          "content": "<p>Interesting idea. !\nI'd considered more \"directed\" downsampling (e.g. keep only IPs in the test set or is_attributed == 1 , but then there's a strong risk of bias/shift, e.g. in the other features) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 304699,
      "author_name": "Skinish",
      "author_url": "",
      "post_date": "2018-03-27T22:35:56.957000",
      "content": "<p>Are you guys training with all the data? If not, how will you choose the subsample? </p>\n\n<p>Moreover, is it worth splitting the dataset in a way that the validation is closer time-wise to the test set? Or is it better to do a stratified separation?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 303078,
      "author_name": "NaN",
      "author_url": "",
      "post_date": "2018-03-25T13:03:33.377000",
      "content": "<p>Thank you for your wonderful ideas but I have yet to manage this dataset.\nOnce I load the dataset for xgboost (DMatrix function), that volume get bigger.\nI think it is useful to save memory to use libsvm-txt format because it can be directly loaded into xgboost but it's difficult to convert.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 299940,
      "author_name": "abhishek kumar",
      "author_url": "",
      "post_date": "2018-03-21T06:15:56.710000",
      "content": "<p>I am just an undergraduate student at iit (ism) , dhanbad , india. How long will it take me to learn everything?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 299946,
          "author_name": "Asparuh Hristov",
          "author_url": "",
          "post_date": "2018-03-21T06:20:53.997000",
          "content": "<p>You can never learn everything in data science... Just do it regualary and enjoy</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 297458,
      "author_name": "Ravi Teja Gutta",
      "author_url": "",
      "post_date": "2018-03-17T04:35:20.330000",
      "content": "<p>Here is another one, we can use out-of-core learning algos(vowpal wabbit), we can use algos that use stochastic gradient descent(logistic regression, neural nets)  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 296499,
      "author_name": "Seymour",
      "author_url": "",
      "post_date": "2018-03-15T10:16:51.537000",
      "content": "<p>is there any packages/approach suggested using R?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 296502,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-03-15T10:18:12.800000",
          "content": "<p>As a starting point, I would say data.table for aggregation / cleanup / processing, after that lightgbm for modeling.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 295970,
      "author_name": "Mutafaf Wahhaj",
      "author_url": "",
      "post_date": "2018-03-14T13:57:50.323000",
      "content": "<p>Good Insights.. I'm new to the field of Data Science, I searched a few datasets those were many hundreds GB (e.g. 800Gb, 900GB etc). How can we work on such datasets when we didn't have that much storage space or comparatively slower Internet connections. What are cloud resources we can utilize with cheaper costs.?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 294650,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-12T11:10:36.890000",
      "content": "<p>\"Count the clicks from each IP and store this information (as a person with experience in online advertisement I can guarantee you that this will be one of the most influential features). Afterwards delete IP.\"</p>\n\n<p>Which will be more useful information:- \n1. calculating this information using both train and test.\n2. calculating separately for test and train</p>\n\n<p>If I do it like the first option, what type of biases it will create?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 294877,
          "author_name": "Asparuh Hristov",
          "author_url": "",
          "post_date": "2018-03-12T18:54:57.180000",
          "content": "<p>I would say number of clicks until conversion</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 294118,
      "author_name": "Nathan Lauga",
      "author_url": "",
      "post_date": "2018-03-11T10:17:27.680000",
      "content": "<p>Thanks for this tips ! </p>\n\n<p>I tried to use dask but when if we use sklearn or XGBoost algorithms, we have to use pandas so it's usefull just for data processing !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 293575,
      "author_name": "zoubairOmar",
      "author_url": "",
      "post_date": "2018-03-10T06:55:14.317000",
      "content": "<p>hello, i'm new to kaggle and datascience, i wonder how i could use this huge volume of data, when i have a poor pc and internet connection. is there a way to run my script in cloud.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 293598,
          "author_name": "tarun",
          "author_url": "",
          "post_date": "2018-03-10T07:47:46.907000",
          "content": "<p>Hey Zoubair, welcome to kaggle, There are kernel that kaggle offers which you could use, it burns cpus held by kaggle &amp; wont put a load on your pc</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 294221,
          "author_name": "zoubairOmar",
          "author_url": "",
          "post_date": "2018-03-11T15:36:14.563000",
          "content": "<p>thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 293308,
      "author_name": "João Pedro Peinado",
      "author_url": "",
      "post_date": "2018-03-09T16:03:51.323000",
      "content": "<p>Thanks for sharing all this information, this could be very helpful to everyone.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 300558,
      "author_name": "Gautham Prabhu",
      "author_url": "",
      "post_date": "2018-03-21T18:03:23.157000",
      "content": "<p>Thanks for the insights</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 296651,
      "author_name": "Wang Yongshuai",
      "author_url": "",
      "post_date": "2018-03-15T15:08:37.663000",
      "content": "<p>wonderful idea, thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 295595,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-13T22:03:45.533000",
      "content": "<p>Thanks! This is useful.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "292758": "Hereby, I have gathered ideas on how to manage the huge amount of data without dropping (or not much) the predictive power. I will start with a very good observation from Konrad Banachewicz:\n\n 1. https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51347\nNamely, get rid of the year and month as they are all the same\n 2. I will propose to get rid of the minutes and seconds as well. One can still detect clicks coming from the same IP within an \"unreasonable\" amount of time by just comparing the date + hour. Even in the beginning simply using only the hour will be more than good enough.\n 3. use dask instead of pandas (way faster!!!!). I have also just read about \"pandas on ray\", have never used it myself but looks extremely promising.\n 4. in my opinion the column \"attributed time\" is totally useless - don't even read it when loading the files (the train data). What does it matter when it was attributed if you have to predict whether it is attributed or not?\n 5. Count the clicks from each IP and store this information (as a person with experience in online advertisement I can guarantee you that this will be one of the most influential features). Afterwards delete IP.\n 6. Actually before deleting it you can also create a column indicating if the user made more than x (say, 5) clicks within an hour.\n 7. Use the correct datatypes when loading the data\n 8. Use gc (garbage collector). Extremely easy to use and very helpful. Constantly delete anything that is not necessary anymore.\n\nPlease, share more ideas/observations if you can think of smt.",
    "294940": "I created a tutorial kernel with some tips on how to manage data.  It includes subsampling, loading in chunks,  creative processing and basic dask tutorial.  \nhttps://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask.   \nHope you guys find it useful!\n\n",
    "295308": "Step 2 i.e. removal of attribute_time column and keeping click_time hr and day. I have done this using python and sqlite3 (inbuilt package of python). Code of the same is here :- https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51823\n\nFeel free to use this. I think after implementing the above stage, data size can be reduced to around 5 Gb which i think is quite manageable. ",
    "303115": "**'the column \"attributed time\" is totally useless'**\nI do not agree with this statement. How to know a priori whether the IPs in which that variable is a date one day after the last click are always reliable?\nThis is only a hypothetical example but I want to point out that it is algorithms and construction techniques that must decide whether a variable is important or not.\nObviously the intuition of the data scientist is important but I think it is good practice to verify the importance of having a variable within the model before eliminating it.",
    "296131": "I do not have enough computing resources to participate in this competition, 16 GB Ram, 16 GB Swap(SSD), even after using data size reduction techniques discussed  in discussions I am not able to complete a full training pass on all data set. <br> I am planing to find a scheme to select a subset of train and validation data from whole train set and use that for model building through out the competition. Any suggestions?",
    "293391": "Thanks for the shout-out - glad you liked the idea  :-)",
    "293073": "I Tried loading this via pyspark ... it does a decent job, Mllib does most of weightlifting of Random forest... ",
    "293009": "good insights. Are u ruling out the possibility of shared IP? I mean, multiple users with same IP can click on the ad without fraudulent intent, ryt?",
    "292783": "Negative downsampling (i.e downsample the majority class). We're looking for the behavior of a small subset, not the general pop (especially with data anonymized). \nhttps://www.kaggle.com/danofer/downsampling-for-fun-speed\n\nI've had excellent results with this approach in RW problems involving more data, albeit non anonymized problems (and with corporate resources/servers handy).",
    "304699": "Are you guys training with all the data? If not, how will you choose the subsample? \n\nMoreover, is it worth splitting the dataset in a way that the validation is closer time-wise to the test set? Or is it better to do a stratified separation?",
    "303078": "Thank you for your wonderful ideas but I have yet to manage this dataset.\nOnce I load the dataset for xgboost (DMatrix function), that volume get bigger.\nI think it is useful to save memory to use libsvm-txt format because it can be directly loaded into xgboost but it's difficult to convert.",
    "299940": "I am just an undergraduate student at iit (ism) , dhanbad , india. How long will it take me to learn everything?",
    "297458": "Here is another one, we can use out-of-core learning algos(vowpal wabbit), we can use algos that use stochastic gradient descent(logistic regression, neural nets)  ",
    "296499": "is there any packages/approach suggested using R?",
    "295970": "Good Insights.. I'm new to the field of Data Science, I searched a few datasets those were many hundreds GB (e.g. 800Gb, 900GB etc). How can we work on such datasets when we didn't have that much storage space or comparatively slower Internet connections. What are cloud resources we can utilize with cheaper costs.?\n\n\n\n",
    "294650": "\"Count the clicks from each IP and store this information (as a person with experience in online advertisement I can guarantee you that this will be one of the most influential features). Afterwards delete IP.\"\n\nWhich will be more useful information:- \n1. calculating this information using both train and test.\n2. calculating separately for test and train\n\nIf I do it like the first option, what type of biases it will create?",
    "294118": "Thanks for this tips ! \n\nI tried to use dask but when if we use sklearn or XGBoost algorithms, we have to use pandas so it's usefull just for data processing !",
    "293575": "hello, i'm new to kaggle and datascience, i wonder how i could use this huge volume of data, when i have a poor pc and internet connection. is there a way to run my script in cloud.",
    "293308": "Thanks for sharing all this information, this could be very helpful to everyone.",
    "300558": "Thanks for the insights",
    "296651": "wonderful idea, thanks!",
    "295595": "Thanks! This is useful."
  }
}