{
  "id": 180056,
  "title": "TF Record Dataset and Own Baseline",
  "url": "/competitions/landmark-recognition-2020/discussion/180056",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2020-09-03T17:26:15.612000",
  "votes": 26,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am training models and it does not take that much time with 60% of the data. Im getting interesting results with a really basic model, (not using arcface, cosine-softmax etc….).</p>\n<p>Using 1 fold validation with 20% of the data give me a validation accuracy of 0.79 and a gap of 0.76.<br>\nThe public leadearboad is 0.1188. Im not using retrieval inference so that could be a reason why the score is so low. Onother point to mention is that my inference pipeline gives the same score as my model validation so i confirm it is correct. At last, the inference pipeline is very fast this way, it predicts the private test set in 30 minutes (because i am using a standard multi classification model, not retrieval approach. Im pretty sure there is a huge terrain of improvement.</p>\n<p>Question to consider: </p>\n<ul>\n<li>The distribution of the target is the same for the training set and the private test set? I believe it is not the same</li>\n<li>How can we resolve the non landmark problem?, we want to predict empty or a really low confidence score for that cases</li>\n</ul>\n<p>Here are the datasets i am using, the images are resize to 384 x 384 and have 100% quality</p>\n<p><a href=\"https://www.kaggle.com/ragnar123/landmark-tfrecords-384\" target=\"_blank\">https://www.kaggle.com/ragnar123/landmark-tfrecords-384</a></p>\n<p><a href=\"https://www.kaggle.com/ragnar123/landmark-tfrecords-384-2\" target=\"_blank\">https://www.kaggle.com/ragnar123/landmark-tfrecords-384-2</a></p>\n<p>Their are 2 because the 50 tf records dont fit in one dataset (more than 100GB). The 50 tf records are stratified by the class and also they are encoded. </p>\n<p>Here is a link where you will find how to map the tf records.</p>\n<p><a href=\"https://www.kaggle.com/ragnar123/landmark-image-train\" target=\"_blank\">https://www.kaggle.com/ragnar123/landmark-image-train</a></p>\n<p>Here their is a column names landmark_id_encode. This is the target label encoded so that you can map the final results of the model to predict the private test.</p>\n<p>Finally, the dataset now has 80% of the data, in the following days i will upload the rest to have the 50 tf records (this is 100% of the training data).</p>\n<p>I hope this is helpfull for training your own models. If you want i can make a notebook with the complete pipeline. Cheers and have fun</p>",
  "messages": [
    {
      "id": 997020,
      "postDate": "2020-09-03T17:26:15.613Z",
      "content": "<p>I am training models and it does not take that much time with 60% of the data. Im getting interesting results with a really basic model, (not using arcface, cosine-softmax etc….).</p>\n<p>Using 1 fold validation with 20% of the data give me a validation accuracy of 0.79 and a gap of 0.76.<br>\nThe public leadearboad is 0.1188. Im not using retrieval inference so that could be a reason why the score is so low. Onother point to mention is that my inference pipeline gives the same score as my model validation so i confirm it is correct. At last, the inference pipeline is very fast this way, it predicts the private test set in 30 minutes (because i am using a standard multi classification model, not retrieval approach. Im pretty sure there is a huge terrain of improvement.</p>\n<p>Question to consider: </p>\n<ul>\n<li>The distribution of the target is the same for the training set and the private test set? I believe it is not the same</li>\n<li>How can we resolve the non landmark problem?, we want to predict empty or a really low confidence score for that cases</li>\n</ul>\n<p>Here are the datasets i am using, the images are resize to 384 x 384 and have 100% quality</p>\n<p><a href=\"https://www.kaggle.com/ragnar123/landmark-tfrecords-384\" target=\"_blank\">https://www.kaggle.com/ragnar123/landmark-tfrecords-384</a></p>\n<p><a href=\"https://www.kaggle.com/ragnar123/landmark-tfrecords-384-2\" target=\"_blank\">https://www.kaggle.com/ragnar123/landmark-tfrecords-384-2</a></p>\n<p>Their are 2 because the 50 tf records dont fit in one dataset (more than 100GB). The 50 tf records are stratified by the class and also they are encoded. </p>\n<p>Here is a link where you will find how to map the tf records.</p>\n<p><a href=\"https://www.kaggle.com/ragnar123/landmark-image-train\" target=\"_blank\">https://www.kaggle.com/ragnar123/landmark-image-train</a></p>\n<p>Here their is a column names landmark_id_encode. This is the target label encoded so that you can map the final results of the model to predict the private test.</p>\n<p>Finally, the dataset now has 80% of the data, in the following days i will upload the rest to have the 50 tf records (this is 100% of the training data).</p>\n<p>I hope this is helpfull for training your own models. If you want i can make a notebook with the complete pipeline. Cheers and have fun</p>",
      "rawMarkdown": "I am training models and it does not take that much time with 60% of the data. Im getting interesting results with a really basic model, (not using arcface, cosine-softmax etc....).\n\nUsing 1 fold validation with 20% of the data give me a validation accuracy of 0.79 and a gap of 0.76.\nThe public leadearboad is 0.1188. Im not using retrieval inference so that could be a reason why the score is so low. Onother point to mention is that my inference pipeline gives the same score as my model validation so i confirm it is correct. At last, the inference pipeline is very fast this way, it predicts the private test set in 30 minutes (because i am using a standard multi classification model, not retrieval approach. Im pretty sure there is a huge terrain of improvement.\n\nQuestion to consider: \n\n- The distribution of the target is the same for the training set and the private test set? I believe it is not the same\n- How can we resolve the non landmark problem?, we want to predict empty or a really low confidence score for that cases\n\nHere are the datasets i am using, the images are resize to 384 x 384 and have 100% quality\n\nhttps://www.kaggle.com/ragnar123/landmark-tfrecords-384\n\nhttps://www.kaggle.com/ragnar123/landmark-tfrecords-384-2\n\nTheir are 2 because the 50 tf records dont fit in one dataset (more than 100GB). The 50 tf records are stratified by the class and also they are encoded. \n\nHere is a link where you will find how to map the tf records.\n\nhttps://www.kaggle.com/ragnar123/landmark-image-train\n\nHere their is a column names landmark_id_encode. This is the target label encoded so that you can map the final results of the model to predict the private test.\n\nFinally, the dataset now has 80% of the data, in the following days i will upload the rest to have the 50 tf records (this is 100% of the training data).\n\nI hope this is helpfull for training your own models. If you want i can make a notebook with the complete pipeline. Cheers and have fun\n\n",
      "votes": 26
    },
    {
      "id": 1013295,
      "postDate": "2020-09-16T16:11:53.673Z",
      "content": "<p>Thank you for the dataset. Unfortunately, I was not able to download this one: <a href=\"https://www.kaggle.com/ragnar123/landmark-tfrecords-384\" target=\"_blank\">https://www.kaggle.com/ragnar123/landmark-tfrecords-384</a> using this Kaggle API command. It returned an error message \"404 - Not Found\". Is it because of storage limitations in your Kaggle dataset?</p>",
      "rawMarkdown": "Thank you for the dataset. Unfortunately, I was not able to download this one: https://www.kaggle.com/ragnar123/landmark-tfrecords-384 using this Kaggle API command. It returned an error message \"404 - Not Found\". Is it because of storage limitations in your Kaggle dataset?"
    },
    {
      "id": 1012062,
      "postDate": "2020-09-15T22:14:10.183Z",
      "content": "<p>Who is there to be team with with me i have more great approach than this</p>",
      "rawMarkdown": "Who is there to be team with with me i have more great approach than this"
    },
    {
      "id": 1012060,
      "postDate": "2020-09-15T22:11:57.213Z",
      "content": "<p>So Helpful<br>\n<a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> </p>",
      "rawMarkdown": "So Helpful\n@ragnar123 "
    },
    {
      "id": 1011138,
      "postDate": "2020-09-15T09:03:38.420Z",
      "content": "<p>I see there is a limit of 20GB on datasets. How were you able to create ~75GB sized datasets?? Thanks</p>",
      "rawMarkdown": "I see there is a limit of 20GB on datasets. How were you able to create ~75GB sized datasets?? Thanks"
    },
    {
      "id": 1005700,
      "postDate": "2020-09-10T17:08:09.040Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1000999,
      "postDate": "2020-09-07T02:07:35.137Z",
      "content": "<p>Thanks, this is very useful！</p>",
      "rawMarkdown": "Thanks, this is very useful！",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1013295,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2020-09-16T16:11:53.673000",
      "content": "<p>Thank you for the dataset. Unfortunately, I was not able to download this one: <a href=\"https://www.kaggle.com/ragnar123/landmark-tfrecords-384\" target=\"_blank\">https://www.kaggle.com/ragnar123/landmark-tfrecords-384</a> using this Kaggle API command. It returned an error message \"404 - Not Found\". Is it because of storage limitations in your Kaggle dataset?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1012062,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-09-15T22:14:10.183000",
      "content": "<p>Who is there to be team with with me i have more great approach than this</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1012060,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2020-09-15T22:11:57.213000",
      "content": "<p>So Helpful<br>\n<a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1011138,
      "author_name": "JkReddy",
      "author_url": "",
      "post_date": "2020-09-15T09:03:38.420000",
      "content": "<p>I see there is a limit of 20GB on datasets. How were you able to create ~75GB sized datasets?? Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1005700,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-10T17:08:09.040000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1000999,
      "author_name": "xckkcxxck",
      "author_url": "",
      "post_date": "2020-09-07T02:07:35.137000",
      "content": "<p>Thanks, this is very useful！</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "997020": "I am training models and it does not take that much time with 60% of the data. Im getting interesting results with a really basic model, (not using arcface, cosine-softmax etc....).\n\nUsing 1 fold validation with 20% of the data give me a validation accuracy of 0.79 and a gap of 0.76.\nThe public leadearboad is 0.1188. Im not using retrieval inference so that could be a reason why the score is so low. Onother point to mention is that my inference pipeline gives the same score as my model validation so i confirm it is correct. At last, the inference pipeline is very fast this way, it predicts the private test set in 30 minutes (because i am using a standard multi classification model, not retrieval approach. Im pretty sure there is a huge terrain of improvement.\n\nQuestion to consider: \n\n- The distribution of the target is the same for the training set and the private test set? I believe it is not the same\n- How can we resolve the non landmark problem?, we want to predict empty or a really low confidence score for that cases\n\nHere are the datasets i am using, the images are resize to 384 x 384 and have 100% quality\n\nhttps://www.kaggle.com/ragnar123/landmark-tfrecords-384\n\nhttps://www.kaggle.com/ragnar123/landmark-tfrecords-384-2\n\nTheir are 2 because the 50 tf records dont fit in one dataset (more than 100GB). The 50 tf records are stratified by the class and also they are encoded. \n\nHere is a link where you will find how to map the tf records.\n\nhttps://www.kaggle.com/ragnar123/landmark-image-train\n\nHere their is a column names landmark_id_encode. This is the target label encoded so that you can map the final results of the model to predict the private test.\n\nFinally, the dataset now has 80% of the data, in the following days i will upload the rest to have the 50 tf records (this is 100% of the training data).\n\nI hope this is helpfull for training your own models. If you want i can make a notebook with the complete pipeline. Cheers and have fun\n\n",
    "1013295": "Thank you for the dataset. Unfortunately, I was not able to download this one: https://www.kaggle.com/ragnar123/landmark-tfrecords-384 using this Kaggle API command. It returned an error message \"404 - Not Found\". Is it because of storage limitations in your Kaggle dataset?",
    "1012062": "Who is there to be team with with me i have more great approach than this",
    "1012060": "So Helpful\n@ragnar123 ",
    "1011138": "I see there is a limit of 20GB on datasets. How were you able to create ~75GB sized datasets?? Thanks",
    "1005700": "",
    "1000999": "Thanks, this is very useful！"
  }
}