{
  "id": 52658,
  "title": "test_supplement.csv",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52658",
  "author_name": "inversion",
  "post_date": "2018-03-21T20:09:38.897000",
  "votes": 43,
  "comment_count": 18,
  "views": 0,
  "content": "<p>A new file has been added to the Data Page and to Kernels: <code>test_supplement.csv</code></p>\n\n<p>This is a larger test set that was unintentionally released at the start of the competition. It is not necessary to use this data, but it is permitted to do so. The official test data is a subset of this data. </p>\n\n<p>We don't plan on supporting this data other than posting it so that everyone has access. </p>",
  "messages": [
    {
      "id": 300665,
      "postDate": "2018-03-21T20:09:38.897Z",
      "content": "<p>A new file has been added to the Data Page and to Kernels: <code>test_supplement.csv</code></p>\n\n<p>This is a larger test set that was unintentionally released at the start of the competition. It is not necessary to use this data, but it is permitted to do so. The official test data is a subset of this data. </p>\n\n<p>We don't plan on supporting this data other than posting it so that everyone has access. </p>",
      "rawMarkdown": "A new file has been added to the Data Page and to Kernels: `test_supplement.csv`\n\nThis is a larger test set that was unintentionally released at the start of the competition. It is not necessary to use this data, but it is permitted to do so. The official test data is a subset of this data. \n\nWe don't plan on supporting this data other than posting it so that everyone has access. ",
      "votes": 43
    },
    {
      "id": 315882,
      "postDate": "2018-04-17T18:11:42.843Z",
      "content": "<p>I wrote an R script to create the click_id mapping between test and test_supplement dataset. The duplicates are removed. Check it out here: <a href=\"https://www.kaggle.com/deepdata/mapping-test-click-id-to-test-supplement-click-id\">https://www.kaggle.com/deepdata/mapping-test-click-id-to-test-supplement-click-id</a>. You can just grab the result mapping file if you do not use R.</p>",
      "rawMarkdown": "I wrote an R script to create the click_id mapping between test and test_supplement dataset. The duplicates are removed. Check it out here: [https://www.kaggle.com/deepdata/mapping-test-click-id-to-test-supplement-click-id][1]. You can just grab the result mapping file if you do not use R.\n\n\n  [1]: https://www.kaggle.com/deepdata/mapping-test-click-id-to-test-supplement-click-id",
      "votes": 3
    },
    {
      "id": 301004,
      "postDate": "2018-03-22T06:54:13.627Z",
      "content": "<p>We need submit predictions of all test data(test_supplement.csv) or subset of test data(test.csv) ?</p>",
      "rawMarkdown": "We need submit predictions of all test data(test_supplement.csv) or subset of test data(test.csv) ?",
      "votes": -2,
      "replies": [
        {
          "id": 302226,
          "postDate": "2018-03-23T20:09:58.127Z",
          "content": "<p>You have to submit the predictions of test.csv not test_supplement.csv . </p>",
          "rawMarkdown": "You have to submit the predictions of test.csv not test_supplement.csv . ",
          "votes": 3
        }
      ]
    },
    {
      "id": 312620,
      "postDate": "2018-04-12T03:50:57.527Z",
      "content": "<p>How do we read such a large dataset in  R can any one help?</p>",
      "rawMarkdown": "How do we read such a large dataset in  R can any one help?",
      "replies": [
        {
          "id": 312624,
          "postDate": "2018-04-12T04:01:34.087Z",
          "content": "<p>Hello Kadarsh, \nI don't think it is really big. Depending on your RAM, you will or not be able to read it. A solution which is not in R is to use DB to do the exploration. There are many that propose a community edition where you have 3 nodes and 1TB for free. I am working with Vertica so I would propose it but you can turn to any DB that proposes this service or you can just use Amazon Web Services which are not really expensive in this case and very helpful.\nBest,\nBadr</p>",
          "rawMarkdown": "Hello Kadarsh, \nI don't think it is really big. Depending on your RAM, you will or not be able to read it. A solution which is not in R is to use DB to do the exploration. There are many that propose a community edition where you have 3 nodes and 1TB for free. I am working with Vertica so I would propose it but you can turn to any DB that proposes this service or you can just use Amazon Web Services which are not really expensive in this case and very helpful.\nBest,\nBadr"
        },
        {
          "id": 312859,
          "postDate": "2018-04-12T13:25:46.317Z",
          "content": "<blockquote>\n  <p><strong>kadarsh51 wrote</strong></p>\n  \n  <blockquote>\n    <p>How do we read such a large dataset in  R can any one help?</p>\n  </blockquote>\n</blockquote>\n\n<p>R with data.table for feature engineering would be super fast and memory efficient in my experience. To avoid hitting memory limit on Kaggle kernel, try processing training data in chunks. </p>\n\n<p>For example, </p>\n\n<pre><code>Process first chunk of 25 million rows, save predictions, clear everything from dir.\nProcess second chunk of 25 million rows, save predictions, clear everything from dir.\n...\nRead all the predictions and do simple/desired average. \n</code></pre>\n\n<p>Check <a href=\"https://www.kaggle.com/pranav84/double-xgb-hist-50-mln-rows\">this script</a>  and log where I'm trying chunk processing as mentioned above. Although it's in Python but hope you can get the idea. btw, I'm not so good with python so please ignore the repetitive codes (or try to create a function). </p>\n\n<p>In R, instead of saving preds in csv/ zip, try saving it to <code>.rds</code>  format. <code>rm(list=ls())</code> would also be useful to clear everything from dir before processing next chunk. </p>",
          "rawMarkdown": "\n&gt; **kadarsh51 wrote**\n&gt; \n&gt; &gt; How do we read such a large dataset in  R can any one help?\n\nR with data.table for feature engineering would be super fast and memory efficient in my experience. To avoid hitting memory limit on Kaggle kernel, try processing training data in chunks. \n\nFor example, \n\n    Process first chunk of 25 million rows, save predictions, clear everything from dir.\n    Process second chunk of 25 million rows, save predictions, clear everything from dir.\n    ...\n    Read all the predictions and do simple/desired average. \n\nCheck [this script][1]  and log where I'm trying chunk processing as mentioned above. Although it's in Python but hope you can get the idea. btw, I'm not so good with python so please ignore the repetitive codes (or try to create a function). \n\nIn R, instead of saving preds in csv/ zip, try saving it to `.rds`  format. `rm(list=ls())` would also be useful to clear everything from dir before processing next chunk. \n\n  [1]: https://www.kaggle.com/pranav84/double-xgb-hist-50-mln-rows",
          "votes": 1
        },
        {
          "id": 312919,
          "postDate": "2018-04-12T14:35:50.387Z",
          "content": "<p>@kadarsh51 see this link <a href=\"https://rpubs.com/msundar/large_data_analysis\">https://rpubs.com/msundar/large_data_analysis</a>.\nHope this helps.</p>",
          "rawMarkdown": "@kadarsh51 see this link https://rpubs.com/msundar/large_data_analysis.\nHope this helps."
        }
      ]
    },
    {
      "id": 300712,
      "postDate": "2018-03-21T20:54:35.203Z",
      "content": "<p>Thank you for clearing out this issue and posting the old data so that everybody has access to it. I fully support this decision as I believe it is the most fair one.</p>",
      "rawMarkdown": "Thank you for clearing out this issue and posting the old data so that everybody has access to it. I fully support this decision as I believe it is the most fair one."
    },
    {
      "id": 309638,
      "postDate": "2018-04-05T17:54:49.823Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 311751,
          "postDate": "2018-04-10T15:54:28.037Z",
          "content": "<p>@inversion said:\n\"We don't plan on supporting this data other than posting it so that everyone has access. \"\nSo, I think they don't plan on supporting this data other than posting it so that everyone has access.\nHowever, you can try comparing the IDs of the records.\nRegards.</p>",
          "rawMarkdown": "@inversion said:\n\"We don't plan on supporting this data other than posting it so that everyone has access. \"\nSo, I think they don't plan on supporting this data other than posting it so that everyone has access.\nHowever, you can try comparing the IDs of the records.\nRegards.",
          "votes": 1
        },
        {
          "id": 311887,
          "postDate": "2018-04-10T20:59:36.670Z",
          "content": "<p>You can use a simple left join in a relational db on click_time1=click_time2 channel1=channel2 etc...</p>",
          "rawMarkdown": "You can use a simple left join in a relational db on click_time1=click_time2 channel1=channel2 etc..."
        },
        {
          "id": 311911,
          "postDate": "2018-04-10T22:17:53.183Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54193\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54193</a></p>\n\n<p>I proposed a solution there</p>",
          "rawMarkdown": "https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54193\n\nI proposed a solution there"
        }
      ]
    },
    {
      "id": 304823,
      "postDate": "2018-03-28T03:58:13.273Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 300952,
      "postDate": "2018-03-22T05:14:14.490Z",
      "content": "<p>thank you </p>",
      "rawMarkdown": "thank you ",
      "votes": 1
    },
    {
      "id": 306489,
      "postDate": "2018-03-30T13:41:16.453Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 303107,
      "postDate": "2018-03-25T14:25:20.793Z",
      "content": "<p>Thanks! It'll be helpful.</p>",
      "rawMarkdown": "Thanks! It'll be helpful."
    },
    {
      "id": 302200,
      "postDate": "2018-03-23T19:33:56.200Z",
      "content": "<p>thank you</p>",
      "rawMarkdown": "thank you"
    },
    {
      "id": 300920,
      "postDate": "2018-03-22T04:10:14.477Z",
      "content": "<p>Great! Thank you!</p>",
      "rawMarkdown": "Great! Thank you!"
    }
  ],
  "comments": [
    {
      "id": 315882,
      "author_name": "DeepData",
      "author_url": "",
      "post_date": "2018-04-17T18:11:42.843000",
      "content": "<p>I wrote an R script to create the click_id mapping between test and test_supplement dataset. The duplicates are removed. Check it out here: <a href=\"https://www.kaggle.com/deepdata/mapping-test-click-id-to-test-supplement-click-id\">https://www.kaggle.com/deepdata/mapping-test-click-id-to-test-supplement-click-id</a>. You can just grab the result mapping file if you do not use R.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 301004,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-03-22T06:54:13.627000",
      "content": "<p>We need submit predictions of all test data(test_supplement.csv) or subset of test data(test.csv) ?</p>",
      "votes": -2,
      "replies": [
        {
          "id": 302226,
          "author_name": "Abhilash Awasthi",
          "author_url": "",
          "post_date": "2018-03-23T20:09:58.127000",
          "content": "<p>You have to submit the predictions of test.csv not test_supplement.csv . </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 312620,
      "author_name": "kadarsh51",
      "author_url": "",
      "post_date": "2018-04-12T03:50:57.527000",
      "content": "<p>How do we read such a large dataset in  R can any one help?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 312624,
          "author_name": "Badr Ouali",
          "author_url": "",
          "post_date": "2018-04-12T04:01:34.087000",
          "content": "<p>Hello Kadarsh, \nI don't think it is really big. Depending on your RAM, you will or not be able to read it. A solution which is not in R is to use DB to do the exploration. There are many that propose a community edition where you have 3 nodes and 1TB for free. I am working with Vertica so I would propose it but you can turn to any DB that proposes this service or you can just use Amazon Web Services which are not really expensive in this case and very helpful.\nBest,\nBadr</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 312859,
          "author_name": "Pranav Pandya",
          "author_url": "",
          "post_date": "2018-04-12T13:25:46.317000",
          "content": "<blockquote>\n  <p><strong>kadarsh51 wrote</strong></p>\n  \n  <blockquote>\n    <p>How do we read such a large dataset in  R can any one help?</p>\n  </blockquote>\n</blockquote>\n\n<p>R with data.table for feature engineering would be super fast and memory efficient in my experience. To avoid hitting memory limit on Kaggle kernel, try processing training data in chunks. </p>\n\n<p>For example, </p>\n\n<pre><code>Process first chunk of 25 million rows, save predictions, clear everything from dir.\nProcess second chunk of 25 million rows, save predictions, clear everything from dir.\n...\nRead all the predictions and do simple/desired average. \n</code></pre>\n\n<p>Check <a href=\"https://www.kaggle.com/pranav84/double-xgb-hist-50-mln-rows\">this script</a>  and log where I'm trying chunk processing as mentioned above. Although it's in Python but hope you can get the idea. btw, I'm not so good with python so please ignore the repetitive codes (or try to create a function). </p>\n\n<p>In R, instead of saving preds in csv/ zip, try saving it to <code>.rds</code>  format. <code>rm(list=ls())</code> would also be useful to clear everything from dir before processing next chunk. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 312919,
          "author_name": "ddrbcn",
          "author_url": "",
          "post_date": "2018-04-12T14:35:50.387000",
          "content": "<p>@kadarsh51 see this link <a href=\"https://rpubs.com/msundar/large_data_analysis\">https://rpubs.com/msundar/large_data_analysis</a>.\nHope this helps.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 300712,
      "author_name": "Asparuh Hristov",
      "author_url": "",
      "post_date": "2018-03-21T20:54:35.203000",
      "content": "<p>Thank you for clearing out this issue and posting the old data so that everybody has access to it. I fully support this decision as I believe it is the most fair one.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 309638,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-05T17:54:49.823000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 311751,
          "author_name": "ddrbcn",
          "author_url": "",
          "post_date": "2018-04-10T15:54:28.037000",
          "content": "<p>@inversion said:\n\"We don't plan on supporting this data other than posting it so that everyone has access. \"\nSo, I think they don't plan on supporting this data other than posting it so that everyone has access.\nHowever, you can try comparing the IDs of the records.\nRegards.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 311887,
          "author_name": "Badr Ouali",
          "author_url": "",
          "post_date": "2018-04-10T20:59:36.670000",
          "content": "<p>You can use a simple left join in a relational db on click_time1=click_time2 channel1=channel2 etc...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 311911,
          "author_name": "Badr Ouali",
          "author_url": "",
          "post_date": "2018-04-10T22:17:53.183000",
          "content": "<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54193\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54193</a></p>\n\n<p>I proposed a solution there</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 304823,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-28T03:58:13.273000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 300952,
      "author_name": "Sean Chow",
      "author_url": "",
      "post_date": "2018-03-22T05:14:14.490000",
      "content": "<p>thank you </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 306489,
      "author_name": "ǝu ɐɯ ıu onɥs ıu",
      "author_url": "",
      "post_date": "2018-03-30T13:41:16.453000",
      "content": "<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 303107,
      "author_name": "kenkoooo",
      "author_url": "",
      "post_date": "2018-03-25T14:25:20.793000",
      "content": "<p>Thanks! It'll be helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 302200,
      "author_name": "Pawas Gupta",
      "author_url": "",
      "post_date": "2018-03-23T19:33:56.200000",
      "content": "<p>thank you</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 300920,
      "author_name": "Muhammad Alfiansyah",
      "author_url": "",
      "post_date": "2018-03-22T04:10:14.477000",
      "content": "<p>Great! Thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "300665": "A new file has been added to the Data Page and to Kernels: `test_supplement.csv`\n\nThis is a larger test set that was unintentionally released at the start of the competition. It is not necessary to use this data, but it is permitted to do so. The official test data is a subset of this data. \n\nWe don't plan on supporting this data other than posting it so that everyone has access. ",
    "315882": "I wrote an R script to create the click_id mapping between test and test_supplement dataset. The duplicates are removed. Check it out here: [https://www.kaggle.com/deepdata/mapping-test-click-id-to-test-supplement-click-id][1]. You can just grab the result mapping file if you do not use R.\n\n\n  [1]: https://www.kaggle.com/deepdata/mapping-test-click-id-to-test-supplement-click-id",
    "301004": "We need submit predictions of all test data(test_supplement.csv) or subset of test data(test.csv) ?",
    "312620": "How do we read such a large dataset in  R can any one help?",
    "300712": "Thank you for clearing out this issue and posting the old data so that everybody has access to it. I fully support this decision as I believe it is the most fair one.",
    "309638": "",
    "304823": "",
    "300952": "thank you ",
    "306489": "Thanks",
    "303107": "Thanks! It'll be helpful.",
    "302200": "thank you",
    "300920": "Great! Thank you!"
  }
}