{
  "id": 52304,
  "title": "Looking for tips to deal diligently with a large dataset",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52304",
  "author_name": "",
  "post_date": "2018-03-18T15:37:41.790429Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>This is the biggest database I've ever worked with. I would like to take this competition as an incentive to improve my skills on SQL  and using GCS APIs. I'm trying to figure out how to do an exploratory analysis of the whole dataset. I state that I am not looking for ad hoc solutions, but for those that can scale with even larger databases.</p>\n\n<p>I had the idea of ​​using GSC's SQL and Datalab APIs, essentially a PostgreSQL or MySQL database to which I would interface with python and a jupyter notebook. But my current knowledge of SQL is elementary while on python I have always and only used pandas for EDA, so I have no idea how much and how SQL is suitable for machine learning tasks, and I do not even know what are the best python libraries to interface with PostgreSQL.</p>\n\n<p>So I would like to ask the data scientists more experienced than me about what solutions are used in industry for problems in the range of 2-100 GB and which are the python libraries I should watch. Thanks.</p>",
  "messages": [
    {
      "id": "298047",
      "postDate": "03/18/2018 15:37:41",
      "content": "<p>This is the biggest database I've ever worked with. I would like to take this competition as an incentive to improve my skills on SQL  and using GCS APIs. I'm trying to figure out how to do an exploratory analysis of the whole dataset. I state that I am not looking for ad hoc solutions, but for those that can scale with even larger databases.</p>\n\n<p>I had the idea of ​​using GSC's SQL and Datalab APIs, essentially a PostgreSQL or MySQL database to which I would interface with python and a jupyter notebook. But my current knowledge of SQL is elementary while on python I have always and only used pandas for EDA, so I have no idea how much and how SQL is suitable for machine learning tasks, and I do not even know what are the best python libraries to interface with PostgreSQL.</p>\n\n<p>So I would like to ask the data scientists more experienced than me about what solutions are used in industry for problems in the range of 2-100 GB and which are the python libraries I should watch. Thanks.</p>",
      "rawMarkdown": "This is the biggest database I've ever worked with. I would like to take this competition as an incentive to improve my skills on SQL  and using GCS APIs. I'm trying to figure out how to do an exploratory analysis of the whole dataset. I state that I am not looking for ad hoc solutions, but for those that can scale with even larger databases.\n\nI had the idea of ​​using GSC's SQL and Datalab APIs, essentially a PostgreSQL or MySQL database to which I would interface with python and a jupyter notebook. But my current knowledge of SQL is elementary while on python I have always and only used pandas for EDA, so I have no idea how much and how SQL is suitable for machine learning tasks, and I do not even know what are the best python libraries to interface with PostgreSQL.\n\nSo I would like to ask the data scientists more experienced than me about what solutions are used in industry for problems in the range of 2-100 GB and which are the python libraries I should watch. Thanks.",
      "votes": null
    },
    {
      "id": "301545",
      "postDate": "03/22/2018 22:25:18",
      "content": "<p>[Total noobie here]</p>\n\n<p>I read the file in a csv and my 8gb ram mac cannot handle the training set.\nI was going through <a href=\"https://www.kaggle.com/getting-started/9512\">this</a>.\nI think I will build my model on subset of data but I don't know if this is a good idea. I am not sure how to 'combine' the models on different datasets.</p>",
      "rawMarkdown": "[Total noobie here]\n\nI read the file in a csv and my 8gb ram mac cannot handle the training set.\nI was going through [this][1].\nI think I will build my model on subset of data but I don't know if this is a good idea. I am not sure how to 'combine' the models on different datasets.\n\n  [1]: https://www.kaggle.com/getting-started/9512",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 301545,
      "author_name": "eternalfool",
      "author_url": "",
      "post_date": "03/22/2018 22:25:18",
      "content": "<p>[Total noobie here]</p>\n\n<p>I read the file in a csv and my 8gb ram mac cannot handle the training set.\nI was going through <a href=\"https://www.kaggle.com/getting-started/9512\">this</a>.\nI think I will build my model on subset of data but I don't know if this is a good idea. I am not sure how to 'combine' the models on different datasets.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "298047": "This is the biggest database I've ever worked with. I would like to take this competition as an incentive to improve my skills on SQL  and using GCS APIs. I'm trying to figure out how to do an exploratory analysis of the whole dataset. I state that I am not looking for ad hoc solutions, but for those that can scale with even larger databases.\n\nI had the idea of ​​using GSC's SQL and Datalab APIs, essentially a PostgreSQL or MySQL database to which I would interface with python and a jupyter notebook. But my current knowledge of SQL is elementary while on python I have always and only used pandas for EDA, so I have no idea how much and how SQL is suitable for machine learning tasks, and I do not even know what are the best python libraries to interface with PostgreSQL.\n\nSo I would like to ask the data scientists more experienced than me about what solutions are used in industry for problems in the range of 2-100 GB and which are the python libraries I should watch. Thanks.",
    "301545": "[Total noobie here]\n\nI read the file in a csv and my 8gb ram mac cannot handle the training set.\nI was going through [this][1].\nI think I will build my model on subset of data but I don't know if this is a good idea. I am not sure how to 'combine' the models on different datasets.\n\n  [1]: https://www.kaggle.com/getting-started/9512"
  },
  "source": "meta"
}