{
  "id": 249965,
  "title": "New to Machine Learning or to Kaggle? Check this out.",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/249965",
  "author_name": "",
  "post_date": "2021-06-30T15:18:22.855789600Z",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n<p>New to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\" target=\"_blank\">how to enter a competition using Kaggle Notebooks</a>.</p>",
  "messages": [
    {
      "id": "1371022",
      "postDate": "06/30/2021 15:18:22",
      "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n<p>New to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\" target=\"_blank\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\" target=\"_blank\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\" target=\"_blank\">how to enter a competition using Kaggle Notebooks</a>.</p>",
      "rawMarkdown": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ).",
      "votes": null
    },
    {
      "id": "1372745",
      "postDate": "07/02/2021 01:38:11",
      "content": "<p>I am new to Kaggle competitions. I do not know what file I need to download and. I see the *.csv file has only two columns, id, and target. Test and train data. Are these image files or *.csv files?</p>\n<p>Could you please help me how to start </p>",
      "rawMarkdown": "I am new to Kaggle competitions. I do not know what file I need to download and. I see the *.csv file has only two columns, id, and target. Test and train data. Are these image files or *.csv files?\n\nCould you please help me how to start",
      "votes": null
    },
    {
      "id": "1374046",
      "postDate": "07/03/2021 01:10:23",
      "content": "<p>I think if you look at the data dictionary,<br>\n<code>sample_submission.csv</code>: this file is the sample file showing format of the final submission file you will upload to kaggle<br>\n<code>training_labels.csv</code>: this contains the target value for each ID in the training dataset</p>\n<p>The folders <code>train</code> and <code>test</code> contain the data for each ID in a .npy file, which in my opinion is the <code>numpy</code> file type. It is more efficient than csv and you can load it faster using the library. You use the training data along with the target value to build your model(s) and make predictions on the test IDs in form of probability that the target exists for that ID. Update this info in the submissions file and done!</p>",
      "rawMarkdown": "I think if you look at the data dictionary,\n`sample_submission.csv`: this file is the sample file showing format of the final submission file you will upload to kaggle\n`training_labels.csv`: this contains the target value for each ID in the training dataset\n\nThe folders `train` and `test` contain the data for each ID in a .npy file, which in my opinion is the `numpy` file type. It is more efficient than csv and you can load it faster using the library. You use the training data along with the target value to build your model(s) and make predictions on the test IDs in form of probability that the target exists for that ID. Update this info in the submissions file and done!",
      "votes": null
    },
    {
      "id": "1374050",
      "postDate": "07/03/2021 01:15:16",
      "content": "<p>Hello all,<br>\nI am new in the field of Data but I am pretty excited to be doing this challenge. I was always looking for an interesting dataset in Physics and this seems like a really good first project to get my foot through the door and experience working on a big project.</p>\n<p>My question is how do you manage/compute such a large dataset (that is fragmented)? What techniques do you employ? Can it be done on a normal laptop or you need to use cloud computing? Do you train over parts of your data in batches or you have to do it all it once?</p>\n<p>(one alternative would be use the kaggle notebooks directly, but I am looking for other solutions if possible)</p>\n<p>Appreciate any and all help/comments! Thanks</p>",
      "rawMarkdown": "Hello all,\nI am new in the field of Data but I am pretty excited to be doing this challenge. I was always looking for an interesting dataset in Physics and this seems like a really good first project to get my foot through the door and experience working on a big project.\n\nMy question is how do you manage/compute such a large dataset (that is fragmented)? What techniques do you employ? Can it be done on a normal laptop or you need to use cloud computing? Do you train over parts of your data in batches or you have to do it all it once?\n\n(one alternative would be use the kaggle notebooks directly, but I am looking for other solutions if possible)\n\nAppreciate any and all help/comments! Thanks",
      "votes": null
    },
    {
      "id": "1375612",
      "postDate": "07/04/2021 10:52:01",
      "content": "<p>I am trying to run the program, but It is taking more than 9 hr. It did not complete.</p>",
      "rawMarkdown": "I am trying to run the program, but It is taking more than 9 hr. It did not complete.",
      "votes": null
    },
    {
      "id": "1383694",
      "postDate": "07/11/2021 06:55:56",
      "content": "<p>bro. did u turn on Gpu ot TPU?</p>",
      "rawMarkdown": "bro. did u turn on Gpu ot TPU?",
      "votes": null
    },
    {
      "id": "1452156",
      "postDate": "08/05/2021 14:31:18",
      "content": "<p>Need Team Member!!<br>\nHi , I am relatively new to the field, looking for a team, Doesn't matter if you are new too. Also if you are looking for someone please take me in I will give my 100%.  Thanks.</p>",
      "rawMarkdown": "Need Team Member!!\nHi , I am relatively new to the field, looking for a team, Doesn't matter if you are new too. Also if you are looking for someone please take me in I will give my 100%.  Thanks.",
      "votes": null
    },
    {
      "id": "1467403",
      "postDate": "08/12/2021 01:01:48",
      "content": "<p>I'm a first time learner and don't know much about it, so I'd like to ask some questions.</p>\n<p>1，What is train_labels? What are they for?</p>\n<p>Does it mean unsupervised learning?<br>\nHow do you know that?<br>\nIs this the correct answer data?</p>\n<p>2，What should I preprocess? I would appreciate it if you could give me a brief explanation verbally and by feeling.</p>\n<p>Don't you all understand?<br>\nFor example, is it possible to reject strange images, and then run the model around?</p>\n<p>I'd like to have a concrete example, just for example.</p>\n<p>I want to win the prize money and go to space, so please help me!</p>",
      "rawMarkdown": "I'm a first time learner and don't know much about it, so I'd like to ask some questions.\n\n1，What is train_labels? What are they for?\n\nDoes it mean unsupervised learning?\nHow do you know that?\nIs this the correct answer data?\n\n2，What should I preprocess? I would appreciate it if you could give me a brief explanation verbally and by feeling.\n\nDon't you all understand?\nFor example, is it possible to reject strange images, and then run the model around?\n\nI'd like to have a concrete example, just for example.\n\nI want to win the prize money and go to space, so please help me!",
      "votes": null
    },
    {
      "id": "1467423",
      "postDate": "08/12/2021 01:21:18",
      "content": "<p>I want to create a test_x for evaluation.<br>\nHow do I make a test_x? Train has labels, but this one doesn't, so I don't know how to make one.</p>",
      "rawMarkdown": "I want to create a test_x for evaluation.\nHow do I make a test_x? Train has labels, but this one doesn't, so I don't know how to make one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1372745,
      "author_name": "rmanoharareddy",
      "author_url": "",
      "post_date": "07/02/2021 01:38:11",
      "content": "<p>I am new to Kaggle competitions. I do not know what file I need to download and. I see the *.csv file has only two columns, id, and target. Test and train data. Are these image files or *.csv files?</p>\n<p>Could you please help me how to start </p>",
      "votes": null,
      "replies": [
        {
          "id": 1374046,
          "author_name": "siddharthpatel45",
          "author_url": "",
          "post_date": "07/03/2021 01:10:23",
          "content": "<p>I think if you look at the data dictionary,<br>\n<code>sample_submission.csv</code>: this file is the sample file showing format of the final submission file you will upload to kaggle<br>\n<code>training_labels.csv</code>: this contains the target value for each ID in the training dataset</p>\n<p>The folders <code>train</code> and <code>test</code> contain the data for each ID in a .npy file, which in my opinion is the <code>numpy</code> file type. It is more efficient than csv and you can load it faster using the library. You use the training data along with the target value to build your model(s) and make predictions on the test IDs in form of probability that the target exists for that ID. Update this info in the submissions file and done!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1374050,
      "author_name": "siddharthpatel45",
      "author_url": "",
      "post_date": "07/03/2021 01:15:16",
      "content": "<p>Hello all,<br>\nI am new in the field of Data but I am pretty excited to be doing this challenge. I was always looking for an interesting dataset in Physics and this seems like a really good first project to get my foot through the door and experience working on a big project.</p>\n<p>My question is how do you manage/compute such a large dataset (that is fragmented)? What techniques do you employ? Can it be done on a normal laptop or you need to use cloud computing? Do you train over parts of your data in batches or you have to do it all it once?</p>\n<p>(one alternative would be use the kaggle notebooks directly, but I am looking for other solutions if possible)</p>\n<p>Appreciate any and all help/comments! Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1375612,
      "author_name": "rmanoharareddy",
      "author_url": "",
      "post_date": "07/04/2021 10:52:01",
      "content": "<p>I am trying to run the program, but It is taking more than 9 hr. It did not complete.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1383694,
          "author_name": "kiranchowdary",
          "author_url": "",
          "post_date": "07/11/2021 06:55:56",
          "content": "<p>bro. did u turn on Gpu ot TPU?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1452156,
      "author_name": "ashutosh2343534",
      "author_url": "",
      "post_date": "08/05/2021 14:31:18",
      "content": "<p>Need Team Member!!<br>\nHi , I am relatively new to the field, looking for a team, Doesn't matter if you are new too. Also if you are looking for someone please take me in I will give my 100%.  Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1467403,
      "author_name": "taiyofuckyou",
      "author_url": "",
      "post_date": "08/12/2021 01:01:48",
      "content": "<p>I'm a first time learner and don't know much about it, so I'd like to ask some questions.</p>\n<p>1，What is train_labels? What are they for?</p>\n<p>Does it mean unsupervised learning?<br>\nHow do you know that?<br>\nIs this the correct answer data?</p>\n<p>2，What should I preprocess? I would appreciate it if you could give me a brief explanation verbally and by feeling.</p>\n<p>Don't you all understand?<br>\nFor example, is it possible to reject strange images, and then run the model around?</p>\n<p>I'd like to have a concrete example, just for example.</p>\n<p>I want to win the prize money and go to space, so please help me!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1467423,
      "author_name": "taiyofuckyou",
      "author_url": "",
      "post_date": "08/12/2021 01:21:18",
      "content": "<p>I want to create a test_x for evaluation.<br>\nHow do I make a test_x? Train has labels, but this one doesn't, so I don't know how to make one.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1371022": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ).",
    "1372745": "I am new to Kaggle competitions. I do not know what file I need to download and. I see the *.csv file has only two columns, id, and target. Test and train data. Are these image files or *.csv files?\n\nCould you please help me how to start",
    "1374046": "I think if you look at the data dictionary,\n`sample_submission.csv`: this file is the sample file showing format of the final submission file you will upload to kaggle\n`training_labels.csv`: this contains the target value for each ID in the training dataset\n\nThe folders `train` and `test` contain the data for each ID in a .npy file, which in my opinion is the `numpy` file type. It is more efficient than csv and you can load it faster using the library. You use the training data along with the target value to build your model(s) and make predictions on the test IDs in form of probability that the target exists for that ID. Update this info in the submissions file and done!",
    "1374050": "Hello all,\nI am new in the field of Data but I am pretty excited to be doing this challenge. I was always looking for an interesting dataset in Physics and this seems like a really good first project to get my foot through the door and experience working on a big project.\n\nMy question is how do you manage/compute such a large dataset (that is fragmented)? What techniques do you employ? Can it be done on a normal laptop or you need to use cloud computing? Do you train over parts of your data in batches or you have to do it all it once?\n\n(one alternative would be use the kaggle notebooks directly, but I am looking for other solutions if possible)\n\nAppreciate any and all help/comments! Thanks",
    "1375612": "I am trying to run the program, but It is taking more than 9 hr. It did not complete.",
    "1383694": "bro. did u turn on Gpu ot TPU?",
    "1452156": "Need Team Member!!\nHi , I am relatively new to the field, looking for a team, Doesn't matter if you are new too. Also if you are looking for someone please take me in I will give my 100%.  Thanks.",
    "1467403": "I'm a first time learner and don't know much about it, so I'd like to ask some questions.\n\n1，What is train_labels? What are they for?\n\nDoes it mean unsupervised learning?\nHow do you know that?\nIs this the correct answer data?\n\n2，What should I preprocess? I would appreciate it if you could give me a brief explanation verbally and by feeling.\n\nDon't you all understand?\nFor example, is it possible to reject strange images, and then run the model around?\n\nI'd like to have a concrete example, just for example.\n\nI want to win the prize money and go to space, so please help me!",
    "1467423": "I want to create a test_x for evaluation.\nHow do I make a test_x? Train has labels, but this one doesn't, so I don't know how to make one."
  },
  "source": "meta"
}