{
  "id": 398138,
  "title": "Clarification on how to use `printer_id`",
  "url": "/competitions/early-detection-of-3d-printing-issues/discussion/398138",
  "author_name": "",
  "post_date": "2023-03-28T18:11:09.736481700Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>This is a question from one of the contestants:</p>\n<blockquote>\n  <p>It states that we cannot use printer id for prediction. But in the description it also says that if model overfits that data then we can use it.</p>\n  <p>So can you elaborate upon this.</p>\n</blockquote>\n<p>My answer:</p>\n<p>It's very easy to train a model that overfits all the printers it has seen in the training dataset. But when such model is used to predict images from an unseen printer, it'll performa poorly. This is bad because the goal of this competition is to train a model that can predict with a good accuracy for any unseen printer (hope this makes intuitive sense).</p>\n<p>Since the model can easily overfits by printer, one possible way (not yet tested. probably won't work) to counter this overfitting is to treat each printer as a \"domain\" and use domain adaption technique. This is what we mean by you can use the <code>printer_id</code> in training.</p>\n<p>However, since the goal of this competition is to train a model that can perform for unseen printers, you can't use the <code>printer_id</code> in the <code>test.csv</code> when you run prediction to generate a submission CSV file. Otherwise, it's not that hard to achieve 100% accuracy and render this competition meaningless. :)</p>\n<p>The same can be said to the <code>print_id</code> field.</p>\n<p>Please comment below if I'm still not clear here.</p>",
  "messages": [
    {
      "id": "2200682",
      "postDate": "03/28/2023 18:11:09",
      "content": "<p>This is a question from one of the contestants:</p>\n<blockquote>\n  <p>It states that we cannot use printer id for prediction. But in the description it also says that if model overfits that data then we can use it.</p>\n  <p>So can you elaborate upon this.</p>\n</blockquote>\n<p>My answer:</p>\n<p>It's very easy to train a model that overfits all the printers it has seen in the training dataset. But when such model is used to predict images from an unseen printer, it'll performa poorly. This is bad because the goal of this competition is to train a model that can predict with a good accuracy for any unseen printer (hope this makes intuitive sense).</p>\n<p>Since the model can easily overfits by printer, one possible way (not yet tested. probably won't work) to counter this overfitting is to treat each printer as a \"domain\" and use domain adaption technique. This is what we mean by you can use the <code>printer_id</code> in training.</p>\n<p>However, since the goal of this competition is to train a model that can perform for unseen printers, you can't use the <code>printer_id</code> in the <code>test.csv</code> when you run prediction to generate a submission CSV file. Otherwise, it's not that hard to achieve 100% accuracy and render this competition meaningless. :)</p>\n<p>The same can be said to the <code>print_id</code> field.</p>\n<p>Please comment below if I'm still not clear here.</p>",
      "rawMarkdown": "This is a question from one of the contestants:\n\n> It states that we cannot use printer id for prediction. But in the description it also says that if model overfits that data then we can use it.\n>\n> So can you elaborate upon this.\n\nMy answer:\n\nIt's very easy to train a model that overfits all the printers it has seen in the training dataset. But when such model is used to predict images from an unseen printer, it'll performa poorly. This is bad because the goal of this competition is to train a model that can predict with a good accuracy for any unseen printer (hope this makes intuitive sense).\n\nSince the model can easily overfits by printer, one possible way (not yet tested. probably won't work) to counter this overfitting is to treat each printer as a \"domain\" and use domain adaption technique. This is what we mean by you can use the `printer_id` in training.\n\nHowever, since the goal of this competition is to train a model that can perform for unseen printers, you can't use the `printer_id` in the `test.csv` when you run prediction to generate a submission CSV file. Otherwise, it's not that hard to achieve 100% accuracy and render this competition meaningless. :)\n\nThe same can be said to the `print_id` field.\n\nPlease comment below if I'm still not clear here.",
      "votes": null
    },
    {
      "id": "2202931",
      "postDate": "03/30/2023 12:05:32",
      "content": "<p>I was focusing on the no extrusion aspect of it all this while since it was mentioned that it is a binary classification task, but I will try this approach as well.</p>",
      "rawMarkdown": "I was focusing on the no extrusion aspect of it all this while since it was mentioned that it is a binary classification task, but I will try this approach as well.",
      "votes": null
    },
    {
      "id": "2202970",
      "postDate": "03/30/2023 12:32:29",
      "content": "<p>The goal of this task is still binary classification. You can (you don't have to) use <code>printer_id</code> and/or <code>print_id</code> in training, but it doesn't change the nature of this task.</p>\n<p>I'm afraid it is still not clear to you what <code>printer_id</code> and <code>print_id</code> are and how you are allowed to use them. Please let me know if this is the case.</p>",
      "rawMarkdown": "The goal of this task is still binary classification. You can (you don't have to) use `printer_id` and/or `print_id` in training, but it doesn't change the nature of this task.\n\nI'm afraid it is still not clear to you what `printer_id` and `print_id` are and how you are allowed to use them. Please let me know if this is the case.",
      "votes": null
    },
    {
      "id": "2206951",
      "postDate": "04/03/2023 03:02:19",
      "content": "<p>print_id thing made sense to me, that there are duplicate values in it, which is present in all other columns as well. Though I am still uncertain how do I use them in the training set?</p>",
      "rawMarkdown": "print_id thing made sense to me, that there are duplicate values in it, which is present in all other columns as well. Though I am still uncertain how do I use them in the training set?",
      "votes": null
    },
    {
      "id": "2207024",
      "postDate": "04/03/2023 04:19:09",
      "content": "<p>Can you elaborate on what you mean by \"duplicate values\"? You mean some inputs are identical? If so, can you give me an example?</p>\n<p>It's up to you how you want to use  printer_id and print_id in training. You can safely ignore them if you don't find them helpful in any ways.</p>",
      "rawMarkdown": "Can you elaborate on what you mean by \"duplicate values\"? You mean some inputs are identical? If so, can you give me an example?\n\nIt's up to you how you want to use  printer_id and print_id in training. You can safely ignore them if you don't find them helpful in any ways.",
      "votes": null
    },
    {
      "id": "2207612",
      "postDate": "04/03/2023 14:17:16",
      "content": "<p>I think there were many entries of similar print_id values in the rows. </p>",
      "rawMarkdown": "I think there were many entries of similar print_id values in the rows.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2202931,
      "author_name": "digvijayyadav",
      "author_url": "",
      "post_date": "03/30/2023 12:05:32",
      "content": "<p>I was focusing on the no extrusion aspect of it all this while since it was mentioned that it is a binary classification task, but I will try this approach as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2202970,
          "author_name": "kennethjiangobico",
          "author_url": "",
          "post_date": "03/30/2023 12:32:29",
          "content": "<p>The goal of this task is still binary classification. You can (you don't have to) use <code>printer_id</code> and/or <code>print_id</code> in training, but it doesn't change the nature of this task.</p>\n<p>I'm afraid it is still not clear to you what <code>printer_id</code> and <code>print_id</code> are and how you are allowed to use them. Please let me know if this is the case.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2206951,
              "author_name": "digvijayyadav",
              "author_url": "",
              "post_date": "04/03/2023 03:02:19",
              "content": "<p>print_id thing made sense to me, that there are duplicate values in it, which is present in all other columns as well. Though I am still uncertain how do I use them in the training set?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2207024,
                  "author_name": "kennethjiangobico",
                  "author_url": "",
                  "post_date": "04/03/2023 04:19:09",
                  "content": "<p>Can you elaborate on what you mean by \"duplicate values\"? You mean some inputs are identical? If so, can you give me an example?</p>\n<p>It's up to you how you want to use  printer_id and print_id in training. You can safely ignore them if you don't find them helpful in any ways.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2207612,
                      "author_name": "digvijayyadav",
                      "author_url": "",
                      "post_date": "04/03/2023 14:17:16",
                      "content": "<p>I think there were many entries of similar print_id values in the rows. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2200682": "This is a question from one of the contestants:\n\n> It states that we cannot use printer id for prediction. But in the description it also says that if model overfits that data then we can use it.\n>\n> So can you elaborate upon this.\n\nMy answer:\n\nIt's very easy to train a model that overfits all the printers it has seen in the training dataset. But when such model is used to predict images from an unseen printer, it'll performa poorly. This is bad because the goal of this competition is to train a model that can predict with a good accuracy for any unseen printer (hope this makes intuitive sense).\n\nSince the model can easily overfits by printer, one possible way (not yet tested. probably won't work) to counter this overfitting is to treat each printer as a \"domain\" and use domain adaption technique. This is what we mean by you can use the `printer_id` in training.\n\nHowever, since the goal of this competition is to train a model that can perform for unseen printers, you can't use the `printer_id` in the `test.csv` when you run prediction to generate a submission CSV file. Otherwise, it's not that hard to achieve 100% accuracy and render this competition meaningless. :)\n\nThe same can be said to the `print_id` field.\n\nPlease comment below if I'm still not clear here.",
    "2202931": "I was focusing on the no extrusion aspect of it all this while since it was mentioned that it is a binary classification task, but I will try this approach as well.",
    "2202970": "The goal of this task is still binary classification. You can (you don't have to) use `printer_id` and/or `print_id` in training, but it doesn't change the nature of this task.\n\nI'm afraid it is still not clear to you what `printer_id` and `print_id` are and how you are allowed to use them. Please let me know if this is the case.",
    "2206951": "print_id thing made sense to me, that there are duplicate values in it, which is present in all other columns as well. Though I am still uncertain how do I use them in the training set?",
    "2207024": "Can you elaborate on what you mean by \"duplicate values\"? You mean some inputs are identical? If so, can you give me an example?\n\nIt's up to you how you want to use  printer_id and print_id in training. You can safely ignore them if you don't find them helpful in any ways.",
    "2207612": "I think there were many entries of similar print_id values in the rows."
  },
  "source": "meta"
}