{
  "id": 174919,
  "title": "Can someone clarify the public and private data set for me.",
  "url": "/competitions/landmark-recognition-2020/discussion/174919",
  "author_name": "thanisornsr",
  "post_date": "2020-08-16T08:14:56.249000",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I have some doubt about the structure of data set and the submitting notebook.</p>\n<p>Let say. If my code read train_data (assume 1000 classes) and test_data. Then I use the one-hot-encoder transform the label in train_data set into one hot. Later, I predict test_data and transform it into label again. When it come to the private data set to run on my code, the whole process repeat again. The private train_data might have 500 classes. Then my trained model predict and transform it into label which is not the same distribution from what it's trained since it's trained on 1000 classes. Am I missing anything? Or it's aimed to train it every time the new data set is loaded?</p>\n<p>Would you mind giving me some suggestion?</p>\n<p>All the best,<br>\nThanisorn</p>",
  "messages": [
    {
      "id": 972086,
      "postDate": "2020-08-16T08:14:56.250Z",
      "content": "<p>Hello,</p>\n<p>I have some doubt about the structure of data set and the submitting notebook.</p>\n<p>Let say. If my code read train_data (assume 1000 classes) and test_data. Then I use the one-hot-encoder transform the label in train_data set into one hot. Later, I predict test_data and transform it into label again. When it come to the private data set to run on my code, the whole process repeat again. The private train_data might have 500 classes. Then my trained model predict and transform it into label which is not the same distribution from what it's trained since it's trained on 1000 classes. Am I missing anything? Or it's aimed to train it every time the new data set is loaded?</p>\n<p>Would you mind giving me some suggestion?</p>\n<p>All the best,<br>\nThanisorn</p>",
      "rawMarkdown": "Hello,\n\nI have some doubt about the structure of data set and the submitting notebook.\n\nLet say. If my code read train_data (assume 1000 classes) and test_data. Then I use the one-hot-encoder transform the label in train_data set into one hot. Later, I predict test_data and transform it into label again. When it come to the private data set to run on my code, the whole process repeat again. The private train_data might have 500 classes. Then my trained model predict and transform it into label which is not the same distribution from what it's trained since it's trained on 1000 classes. Am I missing anything? Or it's aimed to train it every time the new data set is loaded?\n\nWould you mind giving me some suggestion?\n\nAll the best,\nThanisorn",
      "votes": 1
    },
    {
      "id": 972336,
      "postDate": "2020-08-16T13:26:14.637Z",
      "content": "<p>if you want to skip, you can upload your model as dateset and run only inference in the notebook so that it will skip retrain.</p>",
      "rawMarkdown": "if you want to skip, you can upload your model as dateset and run only inference in the notebook so that it will skip retrain."
    },
    {
      "id": 972101,
      "postDate": "2020-08-16T08:33:33.767Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 972336,
      "author_name": "Uday Kumar Gurugubelli",
      "author_url": "",
      "post_date": "2020-08-16T13:26:14.637000",
      "content": "<p>if you want to skip, you can upload your model as dateset and run only inference in the notebook so that it will skip retrain.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 972101,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-16T08:33:33.767000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "972086": "Hello,\n\nI have some doubt about the structure of data set and the submitting notebook.\n\nLet say. If my code read train_data (assume 1000 classes) and test_data. Then I use the one-hot-encoder transform the label in train_data set into one hot. Later, I predict test_data and transform it into label again. When it come to the private data set to run on my code, the whole process repeat again. The private train_data might have 500 classes. Then my trained model predict and transform it into label which is not the same distribution from what it's trained since it's trained on 1000 classes. Am I missing anything? Or it's aimed to train it every time the new data set is loaded?\n\nWould you mind giving me some suggestion?\n\nAll the best,\nThanisorn",
    "972336": "if you want to skip, you can upload your model as dateset and run only inference in the notebook so that it will skip retrain.",
    "972101": ""
  }
}