{
  "id": 311154,
  "title": "How did you learn if datasets size is huge?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/311154",
  "author_name": "",
  "post_date": "2022-03-05T04:26:02.415994400Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi.<br>\nAs Written in tittle, I would like to ask you how to learn if datasets size is very huge.<br>\nIf datasets size is huge, it takes a lot of time to train.<br>\nWhen I want to check a few effects to change parameter a bit, I feel inconvenienced because of taking much time to learn.<br>\nUsually, split datasets??  </p>",
  "messages": [
    {
      "id": "1712574",
      "postDate": "03/05/2022 04:26:02",
      "content": "<p>Hi.<br>\nAs Written in tittle, I would like to ask you how to learn if datasets size is very huge.<br>\nIf datasets size is huge, it takes a lot of time to train.<br>\nWhen I want to check a few effects to change parameter a bit, I feel inconvenienced because of taking much time to learn.<br>\nUsually, split datasets??  </p>",
      "rawMarkdown": "Hi.\nAs Written in tittle, I would like to ask you how to learn if datasets size is very huge.\nIf datasets size is huge, it takes a lot of time to train.\nWhen I want to check a few effects to change parameter a bit, I feel inconvenienced because of taking much time to learn.\nUsually, split datasets??",
      "votes": null
    },
    {
      "id": "1712589",
      "postDate": "03/05/2022 04:49:40",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/haruki741\" target=\"_blank\">@haruki741</a> </p>\n<p>Usually we do Big data while dealing with large datasets with Hadoop or Spark etc. <br>\nBut I assume you're doing with Python for now? In such case, split the datasets - Be caution, do it randomly and don't split on horizontal axis (as that can give you diff results due to various well known possible factors) <br>\nObviously split it into two or if it still doesn't hold up good, make it into three….and your work will be done for sure quite smoothly. </p>\n<p>Happy Kaggling ^_^</p>",
      "rawMarkdown": "Hello @haruki741 \n\nUsually we do Big data while dealing with large datasets with Hadoop or Spark etc. \nBut I assume you're doing with Python for now? In such case, split the datasets - Be caution, do it randomly and don't split on horizontal axis (as that can give you diff results due to various well known possible factors) \nObviously split it into two or if it still doesn't hold up good, make it into three....and your work will be done for sure quite smoothly. \n\nHappy Kaggling ^_^",
      "votes": null
    },
    {
      "id": "1713191",
      "postDate": "03/05/2022 18:25:15",
      "content": "<p>I would suggest you to take a subset of the original data (a subset which resembles the original data) and do the experimentations on it. When you find that a particular experiment work well on the subset of the data, then try it on the full dataset.</p>",
      "rawMarkdown": "I would suggest you to take a subset of the original data (a subset which resembles the original data) and do the experimentations on it. When you find that a particular experiment work well on the subset of the data, then try it on the full dataset.",
      "votes": null
    },
    {
      "id": "1713606",
      "postDate": "03/06/2022 07:33:50",
      "content": "<p>Thanks for the advise!<br>\nso, you mean that I should use the subset of the original data, which is not cropped and resized data in case of this competition?</p>",
      "rawMarkdown": "Thanks for the advise!\nso, you mean that I should use the subset of the original data, which is not cropped and resized data in case of this competition?",
      "votes": null
    },
    {
      "id": "1714358",
      "postDate": "03/06/2022 21:16:36",
      "content": "<p>I think you should start on smaller images to begin with. If you really want some performance at start take a subset of the resized data.</p>",
      "rawMarkdown": "I think you should start on smaller images to begin with. If you really want some performance at start take a subset of the resized data.",
      "votes": null
    },
    {
      "id": "1714451",
      "postDate": "03/07/2022 03:00:13",
      "content": "<p>A  good   GPU is  the most important  thing ,   so  here  is  the key  to deeplearing ------money ,  money  and money   :)</p>",
      "rawMarkdown": "A  good   GPU is  the most important  thing ,   so  here  is  the key  to deeplearing ------money ,  money  and money   :)",
      "votes": null
    },
    {
      "id": "1718574",
      "postDate": "03/11/2022 01:13:14",
      "content": "<p>Not even a good GPU cuts it these days. It's either a better model/algorithm or better/bigger datacenters</p>",
      "rawMarkdown": "Not even a good GPU cuts it these days. It's either a better model/algorithm or better/bigger datacenters",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1712589,
      "author_name": "surajjha101",
      "author_url": "",
      "post_date": "03/05/2022 04:49:40",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/haruki741\" target=\"_blank\">@haruki741</a> </p>\n<p>Usually we do Big data while dealing with large datasets with Hadoop or Spark etc. <br>\nBut I assume you're doing with Python for now? In such case, split the datasets - Be caution, do it randomly and don't split on horizontal axis (as that can give you diff results due to various well known possible factors) <br>\nObviously split it into two or if it still doesn't hold up good, make it into three….and your work will be done for sure quite smoothly. </p>\n<p>Happy Kaggling ^_^</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1713191,
      "author_name": "sadivamadaan",
      "author_url": "",
      "post_date": "03/05/2022 18:25:15",
      "content": "<p>I would suggest you to take a subset of the original data (a subset which resembles the original data) and do the experimentations on it. When you find that a particular experiment work well on the subset of the data, then try it on the full dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1713606,
          "author_name": "haruki741",
          "author_url": "",
          "post_date": "03/06/2022 07:33:50",
          "content": "<p>Thanks for the advise!<br>\nso, you mean that I should use the subset of the original data, which is not cropped and resized data in case of this competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1714358,
          "author_name": "timotheewright",
          "author_url": "",
          "post_date": "03/06/2022 21:16:36",
          "content": "<p>I think you should start on smaller images to begin with. If you really want some performance at start take a subset of the resized data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1714451,
      "author_name": "xujingzhao",
      "author_url": "",
      "post_date": "03/07/2022 03:00:13",
      "content": "<p>A  good   GPU is  the most important  thing ,   so  here  is  the key  to deeplearing ------money ,  money  and money   :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1718574,
          "author_name": "dentistdad",
          "author_url": "",
          "post_date": "03/11/2022 01:13:14",
          "content": "<p>Not even a good GPU cuts it these days. It's either a better model/algorithm or better/bigger datacenters</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1712574": "Hi.\nAs Written in tittle, I would like to ask you how to learn if datasets size is very huge.\nIf datasets size is huge, it takes a lot of time to train.\nWhen I want to check a few effects to change parameter a bit, I feel inconvenienced because of taking much time to learn.\nUsually, split datasets??",
    "1712589": "Hello @haruki741 \n\nUsually we do Big data while dealing with large datasets with Hadoop or Spark etc. \nBut I assume you're doing with Python for now? In such case, split the datasets - Be caution, do it randomly and don't split on horizontal axis (as that can give you diff results due to various well known possible factors) \nObviously split it into two or if it still doesn't hold up good, make it into three....and your work will be done for sure quite smoothly. \n\nHappy Kaggling ^_^",
    "1713191": "I would suggest you to take a subset of the original data (a subset which resembles the original data) and do the experimentations on it. When you find that a particular experiment work well on the subset of the data, then try it on the full dataset.",
    "1713606": "Thanks for the advise!\nso, you mean that I should use the subset of the original data, which is not cropped and resized data in case of this competition?",
    "1714358": "I think you should start on smaller images to begin with. If you really want some performance at start take a subset of the resized data.",
    "1714451": "A  good   GPU is  the most important  thing ,   so  here  is  the key  to deeplearing ------money ,  money  and money   :)",
    "1718574": "Not even a good GPU cuts it these days. It's either a better model/algorithm or better/bigger datacenters"
  },
  "source": "meta"
}