{
  "id": 238933,
  "title": "This competition is perfect for learning data engineering",
  "url": "/competitions/google-smartphone-decimeter-challenge/discussion/238933",
  "author_name": "",
  "post_date": "2021-05-14T02:00:11.529077900Z",
  "votes": 33,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Tabular, text and image datasets are commodities nowadays - you can easily load them with tensorflow, torch, keras, scikit-learn, etc.</p>\n<p>However, geospatial data are more complex by nature, and this particular dataset contains a more complex structure (multiple files per sample, relational structure) than text/tabular/image data. You can't rely on simple wrappers like <code>pd.read_csv</code> or <code>keras.datasets</code> to spoon feed the data. </p>\n<p>Spending time on exploring the {meta, summary, geospatial} data and reading the top EDA notebooks will definitely be beneficial in terms of learning to deal with complex geospatial data.</p>",
  "messages": [
    {
      "id": "1306672",
      "postDate": "05/14/2021 02:00:11",
      "content": "<p>Tabular, text and image datasets are commodities nowadays - you can easily load them with tensorflow, torch, keras, scikit-learn, etc.</p>\n<p>However, geospatial data are more complex by nature, and this particular dataset contains a more complex structure (multiple files per sample, relational structure) than text/tabular/image data. You can't rely on simple wrappers like <code>pd.read_csv</code> or <code>keras.datasets</code> to spoon feed the data. </p>\n<p>Spending time on exploring the {meta, summary, geospatial} data and reading the top EDA notebooks will definitely be beneficial in terms of learning to deal with complex geospatial data.</p>",
      "rawMarkdown": "Tabular, text and image datasets are commodities nowadays - you can easily load them with tensorflow, torch, keras, scikit-learn, etc.\n\nHowever, geospatial data are more complex by nature, and this particular dataset contains a more complex structure (multiple files per sample, relational structure) than text/tabular/image data. You can't rely on simple wrappers like `pd.read_csv` or `keras.datasets` to spoon feed the data. \n\nSpending time on exploring the {meta, summary, geospatial} data and reading the top EDA notebooks will definitely be beneficial in terms of learning to deal with complex geospatial data.",
      "votes": null
    },
    {
      "id": "1307742",
      "postDate": "05/14/2021 16:31:10",
      "content": "<p>You are right. I have been to trying to understand data for 2 days. 😂</p>",
      "rawMarkdown": "You are right. I have been to trying to understand data for 2 days. 😂",
      "votes": null
    },
    {
      "id": "1307788",
      "postDate": "05/14/2021 16:59:35",
      "content": "<p>Takes you 2 days this time, but next time it will take you 2 hours ;)</p>\n<p>Have you taken a look at this thread? <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590\" target=\"_blank\">https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590</a></p>",
      "rawMarkdown": "Takes you 2 days this time, but next time it will take you 2 hours ;)\n\nHave you taken a look at this thread? https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590",
      "votes": null
    },
    {
      "id": "1307833",
      "postDate": "05/14/2021 17:37:26",
      "content": "<p>Yes. It was worth it. <br>\nI've looked at thread and figuring out a way to design a one to one pipeline on this one. :) </p>",
      "rawMarkdown": "Yes. It was worth it. \nI've looked at thread and figuring out a way to design a one to one pipeline on this one. :)",
      "votes": null
    },
    {
      "id": "1357737",
      "postDate": "06/19/2021 23:41:31",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> Agree , the data is sooo overwhleming and requires a lot of domain knowledge to get intuition of various features . </p>",
      "rawMarkdown": "xhlulu Agree , the data is sooo overwhleming and requires a lot of domain knowledge to get intuition of various features .",
      "votes": null
    },
    {
      "id": "1358254",
      "postDate": "06/20/2021 10:01:06",
      "content": "<p>I agree with you!</p>",
      "rawMarkdown": "I agree with you!",
      "votes": null
    },
    {
      "id": "1374146",
      "postDate": "07/03/2021 04:13:38",
      "content": "<p>I agree with that. Im going to check EDA notebooks.</p>",
      "rawMarkdown": "I agree with that. Im going to check EDA notebooks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1307742,
      "author_name": "micheomaano",
      "author_url": "",
      "post_date": "05/14/2021 16:31:10",
      "content": "<p>You are right. I have been to trying to understand data for 2 days. 😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 1307788,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "05/14/2021 16:59:35",
          "content": "<p>Takes you 2 days this time, but next time it will take you 2 hours ;)</p>\n<p>Have you taken a look at this thread? <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590\" target=\"_blank\">https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1307833,
          "author_name": "micheomaano",
          "author_url": "",
          "post_date": "05/14/2021 17:37:26",
          "content": "<p>Yes. It was worth it. <br>\nI've looked at thread and figuring out a way to design a one to one pipeline on this one. :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1357737,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "06/19/2021 23:41:31",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> Agree , the data is sooo overwhleming and requires a lot of domain knowledge to get intuition of various features . </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1358254,
      "author_name": "woosungyoon",
      "author_url": "",
      "post_date": "06/20/2021 10:01:06",
      "content": "<p>I agree with you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1374146,
      "author_name": "tensorchoko",
      "author_url": "",
      "post_date": "07/03/2021 04:13:38",
      "content": "<p>I agree with that. Im going to check EDA notebooks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1306672": "Tabular, text and image datasets are commodities nowadays - you can easily load them with tensorflow, torch, keras, scikit-learn, etc.\n\nHowever, geospatial data are more complex by nature, and this particular dataset contains a more complex structure (multiple files per sample, relational structure) than text/tabular/image data. You can't rely on simple wrappers like `pd.read_csv` or `keras.datasets` to spoon feed the data. \n\nSpending time on exploring the {meta, summary, geospatial} data and reading the top EDA notebooks will definitely be beneficial in terms of learning to deal with complex geospatial data.",
    "1307742": "You are right. I have been to trying to understand data for 2 days. 😂",
    "1307788": "Takes you 2 days this time, but next time it will take you 2 hours ;)\n\nHave you taken a look at this thread? https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/238590",
    "1307833": "Yes. It was worth it. \nI've looked at thread and figuring out a way to design a one to one pipeline on this one. :)",
    "1357737": "xhlulu Agree , the data is sooo overwhleming and requires a lot of domain knowledge to get intuition of various features .",
    "1358254": "I agree with you!",
    "1374146": "I agree with that. Im going to check EDA notebooks."
  },
  "source": "meta"
}