{
  "id": 396986,
  "title": "Do you want more data?",
  "url": "/competitions/early-detection-of-3d-printing-issues/discussion/396986",
  "author_name": "",
  "post_date": "2023-03-23T16:08:31.992461400Z",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>There are about 100k images from 7 printers in the initial data set. We do have more data that we can share. However, we are holding them back because larger data set will require more GPU power to train. We are afraid this may give an unfair advantage to the contestants who have access to more GPU power.</p>\n<p>Please comment below to let us know if you are for or against releasing more data.</p>",
  "messages": [
    {
      "id": "2193957",
      "postDate": "03/23/2023 16:08:31",
      "content": "<p>There are about 100k images from 7 printers in the initial data set. We do have more data that we can share. However, we are holding them back because larger data set will require more GPU power to train. We are afraid this may give an unfair advantage to the contestants who have access to more GPU power.</p>\n<p>Please comment below to let us know if you are for or against releasing more data.</p>",
      "rawMarkdown": "There are about 100k images from 7 printers in the initial data set. We do have more data that we can share. However, we are holding them back because larger data set will require more GPU power to train. We are afraid this may give an unfair advantage to the contestants who have access to more GPU power.\n\nPlease comment below to let us know if you are for or against releasing more data.",
      "votes": null
    },
    {
      "id": "2204543",
      "postDate": "03/31/2023 16:49:42",
      "content": "<p>I feel more data could help avoid overfitting which is one of the main issues while training the model.</p>",
      "rawMarkdown": "I feel more data could help avoid overfitting which is one of the main issues while training the model.",
      "votes": null
    },
    {
      "id": "2204586",
      "postDate": "03/31/2023 17:31:09",
      "content": "<p>Ok now it's one vote in favor of adding more data.</p>\n<p>We can add 7-10 printers. Also there are more prints for existing printers. So we can 2x or even 3x the total amount of data.</p>\n<p>I hope by now most of you have a taste of how much computing power will be required to train the model. If you feel that you are struggling to get enough computing power already, and adding more data will give unfair advantage to other contestants, please speak up here. It's completely ok to speak up against releasing more data. I won't hold it against you.</p>\n<p>I don't want to create situation that is unfair for some contestants. But I don't want to hold back data for no reason if all of you feel more data will help. So please let me know your opinion by commenting below so that we can have a healthy discussion.</p>\n<p>I'll give this issue 3 days to be thoroughly discussed. So if you have any opinion on this issue, make sure you are heard during this period.</p>",
      "rawMarkdown": "Ok now it's one vote in favor of adding more data.\n\nWe can add 7-10 printers. Also there are more prints for existing printers. So we can 2x or even 3x the total amount of data.\n\nI hope by now most of you have a taste of how much computing power will be required to train the model. If you feel that you are struggling to get enough computing power already, and adding more data will give unfair advantage to other contestants, please speak up here. It's completely ok to speak up against releasing more data. I won't hold it against you.\n\nI don't want to create situation that is unfair for some contestants. But I don't want to hold back data for no reason if all of you feel more data will help. So please let me know your opinion by commenting below so that we can have a healthy discussion.\n\nI'll give this issue 3 days to be thoroughly discussed. So if you have any opinion on this issue, make sure you are heard during this period.",
      "votes": null
    },
    {
      "id": "2205459",
      "postDate": "04/01/2023 15:29:19",
      "content": "<blockquote>\n  <p>Ok now it's one vote in favor of adding more data.</p>\n  <p>We can add 7-10 printers. Also there are more prints for existing printers. So we can 2x or even 3x the total amount of data.</p>\n  <p>I hope by now most of you have a taste of how much computing power will be required to train the model. If you feel that you are struggling to get enough computing power already, and adding more data will give unfair advantage to other contestants, please speak up here. It's completely ok to speak up against releasing more data. I won't hold it against you.</p>\n  <p>I don't want to create situation that is unfair for some contestants. But I don't want to hold back data for no reason if all of you feel more data will help. So please let me know your opinion by commenting below so that we can have a healthy discussion.</p>\n  <p>I'll give this issue 3 days to be thoroughly discussed. So if you have any opinion on this issue, make sure you are heard during this period.</p>\n</blockquote>\n<p>I prefer a varied batch of additional data to be supplied, possibly with another CSV file. This varied batch should ideally also include some data from the printers not in the original testing data, so one could possibly check for overfitting and work around it. </p>\n<p>Moreover, are we really requiring a ton of computing power? I'm training my model (&lt; 20 layers) on my RTX 3080 Laptop GPU and finish training in under 10 minutes. This seems like very fast training, mainly because of there aren't many data points. </p>",
      "rawMarkdown": "> Ok now it's one vote in favor of adding more data.\n> \n> We can add 7-10 printers. Also there are more prints for existing printers. So we can 2x or even 3x the total amount of data.\n> \n> I hope by now most of you have a taste of how much computing power will be required to train the model. If you feel that you are struggling to get enough computing power already, and adding more data will give unfair advantage to other contestants, please speak up here. It's completely ok to speak up against releasing more data. I won't hold it against you.\n> \n> I don't want to create situation that is unfair for some contestants. But I don't want to hold back data for no reason if all of you feel more data will help. So please let me know your opinion by commenting below so that we can have a healthy discussion.\n> \n> I'll give this issue 3 days to be thoroughly discussed. So if you have any opinion on this issue, make sure you are heard during this period.\n\nI prefer a varied batch of additional data to be supplied, possibly with another CSV file. This varied batch should ideally also include some data from the printers not in the original testing data, so one could possibly check for overfitting and work around it. \n\nMoreover, are we really requiring a ton of computing power? I'm training my model (< 20 layers) on my RTX 3080 Laptop GPU and finish training in under 10 minutes. This seems like very fast training, mainly because of there aren't many data points.",
      "votes": null
    },
    {
      "id": "2205609",
      "postDate": "04/01/2023 18:11:00",
      "content": "<p>Great. 2 votes in favor of more data now. :)</p>",
      "rawMarkdown": "Great. 2 votes in favor of more data now. :)",
      "votes": null
    },
    {
      "id": "2206059",
      "postDate": "04/02/2023 08:18:14",
      "content": "<p>Adding diversity in the printers would definitely help a lot. Even a very small sampling per printer would still boost things and mage the resulting models more stable (and useful for you!)</p>",
      "rawMarkdown": "Adding diversity in the printers would definitely help a lot. Even a very small sampling per printer would still boost things and mage the resulting models more stable (and useful for you!)",
      "votes": null
    },
    {
      "id": "2206522",
      "postDate": "04/02/2023 16:04:02",
      "content": "<p>I'm already only using a sample of the training data in my approach since I don't own a GPU.  I think we still can all achieve a much higher accuracy with the available data which could encourage more creativity — but I am a bit biased :) </p>",
      "rawMarkdown": "I'm already only using a sample of the training data in my approach since I don't own a GPU.  I think we still can all achieve a much higher accuracy with the available data which could encourage more creativity — but I am a bit biased :)",
      "votes": null
    },
    {
      "id": "2206570",
      "postDate": "04/02/2023 16:55:06",
      "content": "<p>Oh. That's amazing that you're only using a small fraction and getting such amazing results. Quite marvelous. </p>",
      "rawMarkdown": "Oh. That's amazing that you're only using a small fraction and getting such amazing results. Quite marvelous.",
      "votes": null
    },
    {
      "id": "2206710",
      "postDate": "04/02/2023 19:06:59",
      "content": "<p>Thank you everyone for giving us your input publicly or privately! </p>\n<p>Based on your input, we have decided to err on the side of caution against introducing anything that might be unfair for some contestants, hence we will <strong>not</strong> release more data.</p>",
      "rawMarkdown": "Thank you everyone for giving us your input publicly or privately! \n\nBased on your input, we have decided to err on the side of caution against introducing anything that might be unfair for some contestants, hence we will **not** release more data.",
      "votes": null
    },
    {
      "id": "2210469",
      "postDate": "04/05/2023 12:15:04",
      "content": "<p>A subset based on a variety of data could definitely help. I tried training on the entire dataset but it feels like it overfits due to limited number of printers, so with a guess that it could do much better if it generalizes on the printers. I could be wrong though, as I have only been in this challenge for less than a week.</p>",
      "rawMarkdown": "A subset based on a variety of data could definitely help. I tried training on the entire dataset but it feels like it overfits due to limited number of printers, so with a guess that it could do much better if it generalizes on the printers. I could be wrong though, as I have only been in this challenge for less than a week.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2204543,
      "author_name": "suryaaseran",
      "author_url": "",
      "post_date": "03/31/2023 16:49:42",
      "content": "<p>I feel more data could help avoid overfitting which is one of the main issues while training the model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2204586,
      "author_name": "kennethjiangobico",
      "author_url": "",
      "post_date": "03/31/2023 17:31:09",
      "content": "<p>Ok now it's one vote in favor of adding more data.</p>\n<p>We can add 7-10 printers. Also there are more prints for existing printers. So we can 2x or even 3x the total amount of data.</p>\n<p>I hope by now most of you have a taste of how much computing power will be required to train the model. If you feel that you are struggling to get enough computing power already, and adding more data will give unfair advantage to other contestants, please speak up here. It's completely ok to speak up against releasing more data. I won't hold it against you.</p>\n<p>I don't want to create situation that is unfair for some contestants. But I don't want to hold back data for no reason if all of you feel more data will help. So please let me know your opinion by commenting below so that we can have a healthy discussion.</p>\n<p>I'll give this issue 3 days to be thoroughly discussed. So if you have any opinion on this issue, make sure you are heard during this period.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2205459,
          "author_name": "shaheryarsohail",
          "author_url": "",
          "post_date": "04/01/2023 15:29:19",
          "content": "<blockquote>\n  <p>Ok now it's one vote in favor of adding more data.</p>\n  <p>We can add 7-10 printers. Also there are more prints for existing printers. So we can 2x or even 3x the total amount of data.</p>\n  <p>I hope by now most of you have a taste of how much computing power will be required to train the model. If you feel that you are struggling to get enough computing power already, and adding more data will give unfair advantage to other contestants, please speak up here. It's completely ok to speak up against releasing more data. I won't hold it against you.</p>\n  <p>I don't want to create situation that is unfair for some contestants. But I don't want to hold back data for no reason if all of you feel more data will help. So please let me know your opinion by commenting below so that we can have a healthy discussion.</p>\n  <p>I'll give this issue 3 days to be thoroughly discussed. So if you have any opinion on this issue, make sure you are heard during this period.</p>\n</blockquote>\n<p>I prefer a varied batch of additional data to be supplied, possibly with another CSV file. This varied batch should ideally also include some data from the printers not in the original testing data, so one could possibly check for overfitting and work around it. </p>\n<p>Moreover, are we really requiring a ton of computing power? I'm training my model (&lt; 20 layers) on my RTX 3080 Laptop GPU and finish training in under 10 minutes. This seems like very fast training, mainly because of there aren't many data points. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2205609,
              "author_name": "kennethjiangobico",
              "author_url": "",
              "post_date": "04/01/2023 18:11:00",
              "content": "<p>Great. 2 votes in favor of more data now. :)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2206059,
      "author_name": "danofer",
      "author_url": "",
      "post_date": "04/02/2023 08:18:14",
      "content": "<p>Adding diversity in the printers would definitely help a lot. Even a very small sampling per printer would still boost things and mage the resulting models more stable (and useful for you!)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2206522,
      "author_name": "williamshabecoff1",
      "author_url": "",
      "post_date": "04/02/2023 16:04:02",
      "content": "<p>I'm already only using a sample of the training data in my approach since I don't own a GPU.  I think we still can all achieve a much higher accuracy with the available data which could encourage more creativity — but I am a bit biased :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 2206570,
          "author_name": "shaheryarsohail",
          "author_url": "",
          "post_date": "04/02/2023 16:55:06",
          "content": "<p>Oh. That's amazing that you're only using a small fraction and getting such amazing results. Quite marvelous. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2206710,
      "author_name": "kennethjiangobico",
      "author_url": "",
      "post_date": "04/02/2023 19:06:59",
      "content": "<p>Thank you everyone for giving us your input publicly or privately! </p>\n<p>Based on your input, we have decided to err on the side of caution against introducing anything that might be unfair for some contestants, hence we will <strong>not</strong> release more data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2210469,
      "author_name": "hemanthkj",
      "author_url": "",
      "post_date": "04/05/2023 12:15:04",
      "content": "<p>A subset based on a variety of data could definitely help. I tried training on the entire dataset but it feels like it overfits due to limited number of printers, so with a guess that it could do much better if it generalizes on the printers. I could be wrong though, as I have only been in this challenge for less than a week.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2193957": "There are about 100k images from 7 printers in the initial data set. We do have more data that we can share. However, we are holding them back because larger data set will require more GPU power to train. We are afraid this may give an unfair advantage to the contestants who have access to more GPU power.\n\nPlease comment below to let us know if you are for or against releasing more data.",
    "2204543": "I feel more data could help avoid overfitting which is one of the main issues while training the model.",
    "2204586": "Ok now it's one vote in favor of adding more data.\n\nWe can add 7-10 printers. Also there are more prints for existing printers. So we can 2x or even 3x the total amount of data.\n\nI hope by now most of you have a taste of how much computing power will be required to train the model. If you feel that you are struggling to get enough computing power already, and adding more data will give unfair advantage to other contestants, please speak up here. It's completely ok to speak up against releasing more data. I won't hold it against you.\n\nI don't want to create situation that is unfair for some contestants. But I don't want to hold back data for no reason if all of you feel more data will help. So please let me know your opinion by commenting below so that we can have a healthy discussion.\n\nI'll give this issue 3 days to be thoroughly discussed. So if you have any opinion on this issue, make sure you are heard during this period.",
    "2205459": "> Ok now it's one vote in favor of adding more data.\n> \n> We can add 7-10 printers. Also there are more prints for existing printers. So we can 2x or even 3x the total amount of data.\n> \n> I hope by now most of you have a taste of how much computing power will be required to train the model. If you feel that you are struggling to get enough computing power already, and adding more data will give unfair advantage to other contestants, please speak up here. It's completely ok to speak up against releasing more data. I won't hold it against you.\n> \n> I don't want to create situation that is unfair for some contestants. But I don't want to hold back data for no reason if all of you feel more data will help. So please let me know your opinion by commenting below so that we can have a healthy discussion.\n> \n> I'll give this issue 3 days to be thoroughly discussed. So if you have any opinion on this issue, make sure you are heard during this period.\n\nI prefer a varied batch of additional data to be supplied, possibly with another CSV file. This varied batch should ideally also include some data from the printers not in the original testing data, so one could possibly check for overfitting and work around it. \n\nMoreover, are we really requiring a ton of computing power? I'm training my model (< 20 layers) on my RTX 3080 Laptop GPU and finish training in under 10 minutes. This seems like very fast training, mainly because of there aren't many data points.",
    "2205609": "Great. 2 votes in favor of more data now. :)",
    "2206059": "Adding diversity in the printers would definitely help a lot. Even a very small sampling per printer would still boost things and mage the resulting models more stable (and useful for you!)",
    "2206522": "I'm already only using a sample of the training data in my approach since I don't own a GPU.  I think we still can all achieve a much higher accuracy with the available data which could encourage more creativity — but I am a bit biased :)",
    "2206570": "Oh. That's amazing that you're only using a small fraction and getting such amazing results. Quite marvelous.",
    "2206710": "Thank you everyone for giving us your input publicly or privately! \n\nBased on your input, we have decided to err on the side of caution against introducing anything that might be unfair for some contestants, hence we will **not** release more data.",
    "2210469": "A subset based on a variety of data could definitely help. I tried training on the entire dataset but it feels like it overfits due to limited number of printers, so with a guess that it could do much better if it generalizes on the printers. I could be wrong though, as I have only been in this challenge for less than a week."
  },
  "source": "meta"
}