{
  "id": 55121,
  "title": "This competition is too memory demanding",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55121",
  "author_name": "Nomoreaday",
  "post_date": "2018-04-22T11:53:39.959000",
  "votes": 2,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Need more memory to run better models...</p>",
  "messages": [
    {
      "id": 317730,
      "postDate": "2018-04-22T11:53:39.960Z",
      "content": "<p>Need more memory to run better models...</p>",
      "rawMarkdown": "Need more memory to run better models...",
      "votes": 2
    },
    {
      "id": 317981,
      "postDate": "2018-04-23T01:42:39.123Z",
      "content": "<p>it't time to buy some cloud service</p>",
      "rawMarkdown": "it't time to buy some cloud service",
      "replies": [
        {
          "id": 318008,
          "postDate": "2018-04-23T03:16:36.453Z",
          "content": "<p>I perhaps will buy a new computer with enough memory and a good GPU..</p>",
          "rawMarkdown": "I perhaps will buy a new computer with enough memory and a good GPU..",
          "votes": 1
        },
        {
          "id": 318013,
          "postDate": "2018-04-23T03:20:52.533Z",
          "content": "<p>GPU's are overpriced atm due to the crypto craze. Also, any new machine you buy right now will require ddr4 (will not accept ddr3). 64GB of ddr4 alone will cost at last $700... not to talk about CPU cost, mobo cost, GPU cost, etc. Best way to get a decent machine is to camp on consumer2consumer refurb / resell sites. But if you purchase retail in haste, you'll pay for it $$$.</p>",
          "rawMarkdown": "GPU's are overpriced atm due to the crypto craze. Also, any new machine you buy right now will require ddr4 (will not accept ddr3). 64GB of ddr4 alone will cost at last $700... not to talk about CPU cost, mobo cost, GPU cost, etc. Best way to get a decent machine is to camp on consumer2consumer refurb / resell sites. But if you purchase retail in haste, you'll pay for it $$$."
        }
      ]
    },
    {
      "id": 317762,
      "postDate": "2018-04-22T13:48:25.227Z",
      "content": "<p>I think we should be happy they didn't gave us the advertised 3 billion clicks per day they say they collect. ;-)</p>\n\n<p>My noob opinion is that one can and should use batching if memory is a problem. </p>\n\n<p>Pandas supports data reading in chunks see here (<a href=\"https://pandas.pydata.org/pandas-docs/stable/io.html#io-chunking\">https://pandas.pydata.org/pandas-docs/stable/io.html#io-chunking</a>) </p>\n\n<p>For features engineering, with low memory, one should probably import the data in database and aggregate using a database engine. Pandas can read from SQL DB as well.</p>",
      "rawMarkdown": "I think we should be happy they didn't gave us the advertised 3 billion clicks per day they say they collect. ;-)\n\nMy noob opinion is that one can and should use batching if memory is a problem. \n\nPandas supports data reading in chunks see here (https://pandas.pydata.org/pandas-docs/stable/io.html#io-chunking) \n\nFor features engineering, with low memory, one should probably import the data in database and aggregate using a database engine. Pandas can read from SQL DB as well.",
      "replies": [
        {
          "id": 317768,
          "postDate": "2018-04-22T14:09:01.790Z",
          "content": "<p>I tried putting the values in SQL DB using SQLAlchemy in Kaggle Kernel, but the 1GB disk space gets depleted if you read more than 15 M rows.</p>",
          "rawMarkdown": "I tried putting the values in SQL DB using SQLAlchemy in Kaggle Kernel, but the 1GB disk space gets depleted if you read more than 15 M rows."
        },
        {
          "id": 317772,
          "postDate": "2018-04-22T14:19:02.927Z",
          "content": "<p>Sorry, now I understood you want to do everything within the Kaggle Kernel. Have you tried doing the engineering on your own machine and upload the resulted set? </p>\n\n<p>I personally don't think the Kaggle Kernels are useful for training. I don't even think Kaggle (Google) will ever intent to give more resources ... for free ... ;-)</p>",
          "rawMarkdown": "Sorry, now I understood you want to do everything within the Kaggle Kernel. Have you tried doing the engineering on your own machine and upload the resulted set? \n\nI personally don't think the Kaggle Kernels are useful for training. I don't even think Kaggle (Google) will ever intent to give more resources ... for free ... ;-)"
        },
        {
          "id": 318096,
          "postDate": "2018-04-23T07:16:29.683Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 317764,
      "postDate": "2018-04-22T13:53:42.557Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 317766,
          "postDate": "2018-04-22T14:02:55.867Z",
          "content": "<p>this is possible and another possibility is that they overfit to the public lb</p>",
          "rawMarkdown": "this is possible and another possibility is that they overfit to the public lb",
          "votes": 2
        },
        {
          "id": 317767,
          "postDate": "2018-04-22T14:03:04.203Z",
          "content": "<p>Before doing feature selection, I have to compute all features that I have generated. Unfunately this step always goins beyong memory limit. If I could add these features, I am sure that my model can perform better.</p>",
          "rawMarkdown": "Before doing feature selection, I have to compute all features that I have generated. Unfunately this step always goins beyong memory limit. If I could add these features, I am sure that my model can perform better."
        },
        {
          "id": 317770,
          "postDate": "2018-04-22T14:12:47.093Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 317773,
          "postDate": "2018-04-22T14:19:52.337Z",
          "content": "<blockquote>\n  <p>another possibility is that they overfit to the public lb</p>\n</blockquote>\n\n<p>I doubt it for many reasons, including who is on the top, and the number of submissions.</p>",
          "rawMarkdown": "&gt; another possibility is that they overfit to the public lb\n\nI doubt it for many reasons, including who is on the top, and the number of submissions.",
          "votes": -1
        },
        {
          "id": 317774,
          "postDate": "2018-04-22T14:27:25.630Z",
          "content": "<blockquote>\n  <p>I heard Azure also has same policies </p>\n</blockquote>\n\n<p>Yeah ... Azure has same policies  (but just $170 free credit ) </p>\n\n<p>In addition, it is very user friendly ...I didn't need to install anything... All tools I need were already installed on the VM( Anaconda, Jupyter Notebook, Spyder and all commons packages (Xgboost, LGBM, tensorflow, keras, sklearn etc.. ) , R Studio ) ..  And it's 20x faster to upload my submission via the web browser of the VM. </p>",
          "rawMarkdown": "&gt;  I heard Azure also has same policies \n\nYeah ... Azure has same policies  (but just $170 free credit ) \n\nIn addition, it is very user friendly ...I didn't need to install anything... All tools I need were already installed on the VM( Anaconda, Jupyter Notebook, Spyder and all commons packages (Xgboost, LGBM, tensorflow, keras, sklearn etc.. ) , R Studio ) ..  And it's 20x faster to upload my submission via the web browser of the VM. \n\n"
        },
        {
          "id": 317776,
          "postDate": "2018-04-22T14:33:39.653Z",
          "content": "<blockquote>\n  <p>I doubt it for many reasons, including who is on the top, and the number of submissions.</p>\n</blockquote>\n\n<p>(All) the people on the top are using \"few\" features ? </p>",
          "rawMarkdown": "&gt;I doubt it for many reasons, including who is on the top, and the number of submissions.\n\n(All) the people on the top are using \"few\" features ? "
        },
        {
          "id": 317778,
          "postDate": "2018-04-22T14:39:58.813Z",
          "content": "<p>@Serigne, I commented on people at the top overfitting (edited my post to make it clearer)</p>",
          "rawMarkdown": "@Serigne, I commented on people at the top overfitting (edited my post to make it clearer)"
        },
        {
          "id": 317927,
          "postDate": "2018-04-22T21:17:00.830Z",
          "content": "<p>@ZijunYao - do you know of any solid tutorials for setting up GCP for this sort of competition (e.g. spinning up an instance, data storage/transfer, custom stack installation)? I'm pretty comfortable with the  aws suite but have never tried the other dark side.</p>",
          "rawMarkdown": "@ZijunYao - do you know of any solid tutorials for setting up GCP for this sort of competition (e.g. spinning up an instance, data storage/transfer, custom stack installation)? I'm pretty comfortable with the  aws suite but have never tried the other dark side."
        },
        {
          "id": 318123,
          "postDate": "2018-04-23T07:57:22.943Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 317981,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-04-23T01:42:39.123000",
      "content": "<p>it't time to buy some cloud service</p>",
      "votes": 0,
      "replies": [
        {
          "id": 318008,
          "author_name": "Nomoreaday",
          "author_url": "",
          "post_date": "2018-04-23T03:16:36.453000",
          "content": "<p>I perhaps will buy a new computer with enough memory and a good GPU..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 318013,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-23T03:20:52.533000",
          "content": "<p>GPU's are overpriced atm due to the crypto craze. Also, any new machine you buy right now will require ddr4 (will not accept ddr3). 64GB of ddr4 alone will cost at last $700... not to talk about CPU cost, mobo cost, GPU cost, etc. Best way to get a decent machine is to camp on consumer2consumer refurb / resell sites. But if you purchase retail in haste, you'll pay for it $$$.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 317762,
      "author_name": "Mihai Cvasnievschi",
      "author_url": "",
      "post_date": "2018-04-22T13:48:25.227000",
      "content": "<p>I think we should be happy they didn't gave us the advertised 3 billion clicks per day they say they collect. ;-)</p>\n\n<p>My noob opinion is that one can and should use batching if memory is a problem. </p>\n\n<p>Pandas supports data reading in chunks see here (<a href=\"https://pandas.pydata.org/pandas-docs/stable/io.html#io-chunking\">https://pandas.pydata.org/pandas-docs/stable/io.html#io-chunking</a>) </p>\n\n<p>For features engineering, with low memory, one should probably import the data in database and aggregate using a database engine. Pandas can read from SQL DB as well.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 317768,
          "author_name": "Sudeep Shukla",
          "author_url": "",
          "post_date": "2018-04-22T14:09:01.790000",
          "content": "<p>I tried putting the values in SQL DB using SQLAlchemy in Kaggle Kernel, but the 1GB disk space gets depleted if you read more than 15 M rows.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317772,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-04-22T14:19:02.927000",
          "content": "<p>Sorry, now I understood you want to do everything within the Kaggle Kernel. Have you tried doing the engineering on your own machine and upload the resulted set? </p>\n\n<p>I personally don't think the Kaggle Kernels are useful for training. I don't even think Kaggle (Google) will ever intent to give more resources ... for free ... ;-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318096,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-23T07:16:29.683000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 317764,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-22T13:53:42.557000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 317766,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-22T14:02:55.867000",
          "content": "<p>this is possible and another possibility is that they overfit to the public lb</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 317767,
          "author_name": "Nomoreaday",
          "author_url": "",
          "post_date": "2018-04-22T14:03:04.203000",
          "content": "<p>Before doing feature selection, I have to compute all features that I have generated. Unfunately this step always goins beyong memory limit. If I could add these features, I am sure that my model can perform better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317770,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-22T14:12:47.093000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317773,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-22T14:19:52.337000",
          "content": "<blockquote>\n  <p>another possibility is that they overfit to the public lb</p>\n</blockquote>\n\n<p>I doubt it for many reasons, including who is on the top, and the number of submissions.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 317774,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-04-22T14:27:25.630000",
          "content": "<blockquote>\n  <p>I heard Azure also has same policies </p>\n</blockquote>\n\n<p>Yeah ... Azure has same policies  (but just $170 free credit ) </p>\n\n<p>In addition, it is very user friendly ...I didn't need to install anything... All tools I need were already installed on the VM( Anaconda, Jupyter Notebook, Spyder and all commons packages (Xgboost, LGBM, tensorflow, keras, sklearn etc.. ) , R Studio ) ..  And it's 20x faster to upload my submission via the web browser of the VM. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317776,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2018-04-22T14:33:39.653000",
          "content": "<blockquote>\n  <p>I doubt it for many reasons, including who is on the top, and the number of submissions.</p>\n</blockquote>\n\n<p>(All) the people on the top are using \"few\" features ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317778,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-22T14:39:58.813000",
          "content": "<p>@Serigne, I commented on people at the top overfitting (edited my post to make it clearer)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 317927,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-04-22T21:17:00.830000",
          "content": "<p>@ZijunYao - do you know of any solid tutorials for setting up GCP for this sort of competition (e.g. spinning up an instance, data storage/transfer, custom stack installation)? I'm pretty comfortable with the  aws suite but have never tried the other dark side.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 318123,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-23T07:57:22.943000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "317730": "Need more memory to run better models...",
    "317981": "it't time to buy some cloud service",
    "317762": "I think we should be happy they didn't gave us the advertised 3 billion clicks per day they say they collect. ;-)\n\nMy noob opinion is that one can and should use batching if memory is a problem. \n\nPandas supports data reading in chunks see here (https://pandas.pydata.org/pandas-docs/stable/io.html#io-chunking) \n\nFor features engineering, with low memory, one should probably import the data in database and aggregate using a database engine. Pandas can read from SQL DB as well.",
    "317764": ""
  }
}