{
  "id": 495267,
  "title": "New to Data Science few quick questions",
  "url": "/competitions/leash-BELKA/discussion/495267",
  "author_name": "",
  "post_date": "2024-04-20T12:18:20.116171800Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, I'm new to data science competitions but am really enjoying them and was wondering if someone can help me out with a few basic questions:</p>\n<p>1:  How do I work with such a large dataset?  It won't even fit on my computer when I unzip it.  I was thinking of buying a usb storage if that would work?  Can I upload something this large to google collab?  How do I do my initial data inspection?</p>\n<p>2:  How do I do my training and testing split, making sure my data sets are as accurate and unbiased as possible?  I'm not asking for all of your secrets, but if you can point me in the right direction I would really appreciate it. </p>\n<p>Thanks so much everyone!<br>\n-Matt</p>",
  "messages": [
    {
      "id": "2763241",
      "postDate": "04/20/2024 12:18:20",
      "content": "<p>Hi, I'm new to data science competitions but am really enjoying them and was wondering if someone can help me out with a few basic questions:</p>\n<p>1:  How do I work with such a large dataset?  It won't even fit on my computer when I unzip it.  I was thinking of buying a usb storage if that would work?  Can I upload something this large to google collab?  How do I do my initial data inspection?</p>\n<p>2:  How do I do my training and testing split, making sure my data sets are as accurate and unbiased as possible?  I'm not asking for all of your secrets, but if you can point me in the right direction I would really appreciate it. </p>\n<p>Thanks so much everyone!<br>\n-Matt</p>",
      "rawMarkdown": "Hi, I'm new to data science competitions but am really enjoying them and was wondering if someone can help me out with a few basic questions:\n\n1:  How do I work with such a large dataset?  It won't even fit on my computer when I unzip it.  I was thinking of buying a usb storage if that would work?  Can I upload something this large to google collab?  How do I do my initial data inspection?\n\n2:  How do I do my training and testing split, making sure my data sets are as accurate and unbiased as possible?  I'm not asking for all of your secrets, but if you can point me in the right direction I would really appreciate it. \n\nThanks so much everyone!\n-Matt",
      "votes": null
    },
    {
      "id": "2763894",
      "postDate": "04/20/2024 19:25:51",
      "content": "<p>I highly recommend using duckdb (see leash tutorial notebooks) to start off with a subset of the data and go from there. That said, <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> has a great notebook with a condensed version of the train and test sets that fit nicely in memory.</p>\n<p>Given that there are some large differences in some of the test set from the train set, I would recommend a scaffold splitting strategy that either usus building blocks or cheminformatics methods like Murcko decomposition. Note: Murcko decomposition can be done with RDKit, a package that is also used in the leash tutorial notebooks.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "I highly recommend using duckdb (see leash tutorial notebooks) to start off with a subset of the data and go from there. That said, @shlomoron has a great notebook with a condensed version of the train and test sets that fit nicely in memory.\n\nGiven that there are some large differences in some of the test set from the train set, I would recommend a scaffold splitting strategy that either usus building blocks or cheminformatics methods like Murcko decomposition. Note: Murcko decomposition can be done with RDKit, a package that is also used in the leash tutorial notebooks.\n\nGood luck!",
      "votes": null
    },
    {
      "id": "2763906",
      "postDate": "04/20/2024 19:35:05",
      "content": "<p>Thanks so much!  Happy hacking!</p>",
      "rawMarkdown": "Thanks so much!  Happy hacking!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2763894,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "04/20/2024 19:25:51",
      "content": "<p>I highly recommend using duckdb (see leash tutorial notebooks) to start off with a subset of the data and go from there. That said, <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> has a great notebook with a condensed version of the train and test sets that fit nicely in memory.</p>\n<p>Given that there are some large differences in some of the test set from the train set, I would recommend a scaffold splitting strategy that either usus building blocks or cheminformatics methods like Murcko decomposition. Note: Murcko decomposition can be done with RDKit, a package that is also used in the leash tutorial notebooks.</p>\n<p>Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2763906,
          "author_name": "mgoral1",
          "author_url": "",
          "post_date": "04/20/2024 19:35:05",
          "content": "<p>Thanks so much!  Happy hacking!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2763241": "Hi, I'm new to data science competitions but am really enjoying them and was wondering if someone can help me out with a few basic questions:\n\n1:  How do I work with such a large dataset?  It won't even fit on my computer when I unzip it.  I was thinking of buying a usb storage if that would work?  Can I upload something this large to google collab?  How do I do my initial data inspection?\n\n2:  How do I do my training and testing split, making sure my data sets are as accurate and unbiased as possible?  I'm not asking for all of your secrets, but if you can point me in the right direction I would really appreciate it. \n\nThanks so much everyone!\n-Matt",
    "2763894": "I highly recommend using duckdb (see leash tutorial notebooks) to start off with a subset of the data and go from there. That said, @shlomoron has a great notebook with a condensed version of the train and test sets that fit nicely in memory.\n\nGiven that there are some large differences in some of the test set from the train set, I would recommend a scaffold splitting strategy that either usus building blocks or cheminformatics methods like Murcko decomposition. Note: Murcko decomposition can be done with RDKit, a package that is also used in the leash tutorial notebooks.\n\nGood luck!",
    "2763906": "Thanks so much!  Happy hacking!"
  },
  "source": "meta"
}