{
  "id": 163350,
  "title": "Your Second Submission (model training/exporting example)",
  "url": "/competitions/landmark-retrieval-2020/discussion/163350",
  "author_name": "Andre Araujo",
  "post_date": "2020-07-01T17:58:04.315000",
  "votes": 42,
  "comment_count": 18,
  "views": 0,
  "content": "<p>After trying out <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163341\">your first submission</a>, now it’s time to train your own landmark retrieval model!</p>\n\n<p>You could use the <a href=\"https://github.com/tensorflow/models/tree/master/research/delf/\">DELF TensorFlow model training codebase</a>. To train your first ResNet50-based model, you can follow the step-by-step instructions outlined <a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training\">here</a>, which cover in detail:</p>\n\n<ul>\n<li>codebase installation,</li>\n<li>GLDv2 image downloading,</li>\n<li>training data preparation, </li>\n<li>and (finally!) model training.</li>\n</ul>\n\n<p>Once model training is complete, you can use the <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py\">export_global_model</a> script to produce a model in the exact format required by the competition. See the command mentioned <a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#kaggle-compatible-global-feature-model\">here</a> which shows how flags should be set.</p>\n\n<p>With this exported model, you can then follow the instructions in the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163341\">first submission post</a> to upload the model into Kaggle for scoring.</p>\n\n<p>We hope this provides you a good place to start, with references for model training and concrete examples for model exporting in the right format.</p>\n\n<p>Cheers,</p>\n\n<p>Andre</p>\n\n<p>PS1: One small note about the above-mentioned training code example: the default training process is set up to train a model with DELF local features (which is an additional head on top of the ResNet backbone). For the purposes of this year’s landmark retrieval challenge, this local feature head is not required (but it’s also not doing any harm right now). We plan on updating the codebase later with training examples where only the global feature head is present.</p>\n\n<p>PS2: Please note that, while these instructions should work and provide a decent model, they are not currently tuned for optimal performance in this challenge (this codebase is work-in-progress). In particular, right now we observe lower performance using these instructions, compared to when using the above-mentioned first submission (which was trained using a different codebase). In the coming weeks, we will continue to add improvements to the codebase, such as: GeM pooling, ArcFace loss, etc -- which will certainly help improve performance a lot. We also welcome contributions from the community, so feel free to send pull requests!</p>",
  "messages": [
    {
      "id": 911411,
      "postDate": "2020-07-01T17:58:04.317Z",
      "content": "<p>After trying out <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163341\">your first submission</a>, now it’s time to train your own landmark retrieval model!</p>\n\n<p>You could use the <a href=\"https://github.com/tensorflow/models/tree/master/research/delf/\">DELF TensorFlow model training codebase</a>. To train your first ResNet50-based model, you can follow the step-by-step instructions outlined <a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training\">here</a>, which cover in detail:</p>\n\n<ul>\n<li>codebase installation,</li>\n<li>GLDv2 image downloading,</li>\n<li>training data preparation, </li>\n<li>and (finally!) model training.</li>\n</ul>\n\n<p>Once model training is complete, you can use the <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py\">export_global_model</a> script to produce a model in the exact format required by the competition. See the command mentioned <a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#kaggle-compatible-global-feature-model\">here</a> which shows how flags should be set.</p>\n\n<p>With this exported model, you can then follow the instructions in the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163341\">first submission post</a> to upload the model into Kaggle for scoring.</p>\n\n<p>We hope this provides you a good place to start, with references for model training and concrete examples for model exporting in the right format.</p>\n\n<p>Cheers,</p>\n\n<p>Andre</p>\n\n<p>PS1: One small note about the above-mentioned training code example: the default training process is set up to train a model with DELF local features (which is an additional head on top of the ResNet backbone). For the purposes of this year’s landmark retrieval challenge, this local feature head is not required (but it’s also not doing any harm right now). We plan on updating the codebase later with training examples where only the global feature head is present.</p>\n\n<p>PS2: Please note that, while these instructions should work and provide a decent model, they are not currently tuned for optimal performance in this challenge (this codebase is work-in-progress). In particular, right now we observe lower performance using these instructions, compared to when using the above-mentioned first submission (which was trained using a different codebase). In the coming weeks, we will continue to add improvements to the codebase, such as: GeM pooling, ArcFace loss, etc -- which will certainly help improve performance a lot. We also welcome contributions from the community, so feel free to send pull requests!</p>",
      "rawMarkdown": "After trying out [your first submission](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163341), now it’s time to train your own landmark retrieval model!\n\nYou could use the [DELF TensorFlow model training codebase](https://github.com/tensorflow/models/tree/master/research/delf/). To train your first ResNet50-based model, you can follow the step-by-step instructions outlined [here](https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training), which cover in detail:\n\n* codebase installation,\n* GLDv2 image downloading,\n* training data preparation, \n* and (finally!) model training.\n\nOnce model training is complete, you can use the [export\\_global\\_model](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py) script to produce a model in the exact format required by the competition. See the command mentioned [here](https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#kaggle-compatible-global-feature-model) which shows how flags should be set.\n\nWith this exported model, you can then follow the instructions in the [first submission post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163341) to upload the model into Kaggle for scoring.\n\nWe hope this provides you a good place to start, with references for model training and concrete examples for model exporting in the right format.\n\nCheers,\n\nAndre\n\nPS1: One small note about the above-mentioned training code example: the default training process is set up to train a model with DELF local features (which is an additional head on top of the ResNet backbone). For the purposes of this year’s landmark retrieval challenge, this local feature head is not required (but it’s also not doing any harm right now). We plan on updating the codebase later with training examples where only the global feature head is present.\n\nPS2: Please note that, while these instructions should work and provide a decent model, they are not currently tuned for optimal performance in this challenge (this codebase is work-in-progress). In particular, right now we observe lower performance using these instructions, compared to when using the above-mentioned first submission (which was trained using a different codebase). In the coming weeks, we will continue to add improvements to the codebase, such as: GeM pooling, ArcFace loss, etc -- which will certainly help improve performance a lot. We also welcome contributions from the community, so feel free to send pull requests!",
      "votes": 41
    },
    {
      "id": 941712,
      "postDate": "2020-07-23T11:18:15.193Z",
      "content": "<p>Hi, thanks for sharing this. A little bit confused about preparing data for training. When generating tfrecord files, should I leave arg <code>train_clean_csv_path</code>empty if data downloading from kaggle(clean data).</p>",
      "rawMarkdown": "Hi, thanks for sharing this. A little bit confused about preparing data for training. When generating tfrecord files, should I leave arg `train_clean_csv_path`empty if data downloading from kaggle(clean data).",
      "votes": 1
    },
    {
      "id": 931214,
      "postDate": "2020-07-16T04:15:56.337Z",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> It seems that we can't restart training form saved_model.pb (I want to try fine-tuning with the baseline model). Could you share the trained ckpt model which you set to run <code>export_global_model.py</code> ?</p>",
      "rawMarkdown": "@andrefaraujo It seems that we can't restart training form saved_model.pb (I want to try fine-tuning with the baseline model). Could you share the trained ckpt model which you set to run `export_global_model.py` ?",
      "votes": 2,
      "replies": [
        {
          "id": 932030,
          "postDate": "2020-07-16T16:27:49.963Z",
          "content": "<p>I can't release the checkpoint.</p>\n\n<p>The baseline model was trained with TF-slim, so one idea would be to convert the ResNet101 weights from TF-slim-resnet to Keras-resnet, by looping over the variables and associating the corresponding ones. You then would set the Keras weights directly, using set_weights -- similar to <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/resnet50.py#L421\">this example</a> where we load ImageNet weights. I haven't done it before, but I would expect it to be doable.</p>\n\n<p>Hope this helps!</p>",
          "rawMarkdown": "I can't release the checkpoint.\n\nThe baseline model was trained with TF-slim, so one idea would be to convert the ResNet101 weights from TF-slim-resnet to Keras-resnet, by looping over the variables and associating the corresponding ones. You then would set the Keras weights directly, using set_weights -- similar to [this example](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/resnet50.py#L421) where we load ImageNet weights. I haven't done it before, but I would expect it to be doable.\n\nHope this helps!",
          "votes": -1
        },
        {
          "id": 942468,
          "postDate": "2020-07-23T18:56:12.173Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 945422,
          "postDate": "2020-07-25T19:48:27.080Z",
          "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> loaded model doesn't have reference to its variables. In your example, it is loaded from h5 file.</p>",
          "rawMarkdown": "@andrefaraujo loaded model doesn't have reference to its variables. In your example, it is loaded from h5 file.",
          "votes": 1
        }
      ]
    },
    {
      "id": 919918,
      "postDate": "2020-07-08T07:29:57.377Z",
      "content": "<p>Hi <a href=\"/andrefaraujo\">@andrefaraujo</a>,</p>\n\n<p>Can we apply your instructions from above and run it in a nb within kaggle? It's not obvious if it we can and my recent attempts point to the answer \"no\".</p>\n\n<p>Are you expecting us to use GCP for downloading data and training?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi @andrefaraujo,\n\nCan we apply your instructions from above and run it in a nb within kaggle? It's not obvious if it we can and my recent attempts point to the answer \"no\".\n\nAre you expecting us to use GCP for downloading data and training?\n\nThanks\n",
      "votes": 2
    },
    {
      "id": 947985,
      "postDate": "2020-07-27T15:37:46.893Z",
      "content": "<p>Hello, does anyone know how to make sure that every .tar file actually unpacks into a folder?  I run download_dataset.sh and some of the files don't check out with their .md5 checksums and exit.  Does anyone know how to fix this?</p>",
      "rawMarkdown": "Hello, does anyone know how to make sure that every .tar file actually unpacks into a folder?  I run download_dataset.sh and some of the files don't check out with their .md5 checksums and exit.  Does anyone know how to fix this?"
    },
    {
      "id": 942117,
      "postDate": "2020-07-23T15:35:45.783Z",
      "content": "<p>Hi <a href=\"/andrefaraujo\">@andrefaraujo</a> <a href=\"/bingyicao\">@bingyicao</a> <a href=\"/camaskew\">@camaskew</a>   , I went through the paper related to DELG. I'm hoping for a little clarity on the training procedure , you guys adopted. Do I understand correct, that R101 was trained only till convolution block  5, using Attention loss and reconstruction loss, along with ArcFace loss? Was ArcFace Loss used to train all the rest of the layers or only convolution block 4?</p>",
      "rawMarkdown": "Hi @andrefaraujo @bingyicao @camaskew   , I went through the paper related to DELG. I'm hoping for a little clarity on the training procedure , you guys adopted. Do I understand correct, that R101 was trained only till convolution block  5, using Attention loss and reconstruction loss, along with ArcFace loss? Was ArcFace Loss used to train all the rest of the layers or only convolution block 4?",
      "replies": [
        {
          "id": 942585,
          "postDate": "2020-07-23T20:58:05.137Z",
          "content": "<p>Hey Leo! For the DELG training is end-to-end using training data with image-level label. Yes, you are correct. During the training, we applied attention_loss, reconstruction_loss and ArcFace_loss to optimize the attention network, AutoEncoder and the classifier correspondingly. ArcFace Loss is a loss function modification, so it's not really applied to layers. Let us know if you have further questions!</p>",
          "rawMarkdown": "Hey Leo! For the DELG training is end-to-end using training data with image-level label. Yes, you are correct. During the training, we applied attention_loss, reconstruction_loss and ArcFace_loss to optimize the attention network, AutoEncoder and the classifier correspondingly. ArcFace Loss is a loss function modification, so it's not really applied to layers. Let us know if you have further questions!",
          "votes": 2
        }
      ]
    },
    {
      "id": 934170,
      "postDate": "2020-07-18T09:22:21.917Z",
      "content": "<p>Thanks for sharing DELF method. Could you show the mAP scores using DELF global model？ I  get 12.7 mAP and I am not sure whether I trained the model correctly.</p>",
      "rawMarkdown": "Thanks for sharing DELF method. Could you show the mAP scores using DELF global model？ I  get 12.7 mAP and I am not sure whether I trained the model correctly."
    },
    {
      "id": 920672,
      "postDate": "2020-07-08T18:42:25.843Z",
      "content": "<p>Hi <a href=\"/andrefaraujo\">@andrefaraujo</a> , \nIs there a timeline for when the other challenges in this workshop going to come out?\nThanks</p>",
      "rawMarkdown": "Hi @andrefaraujo , \nIs there a timeline for when the other challenges in this workshop going to come out?\nThanks"
    },
    {
      "id": 915994,
      "postDate": "2020-07-05T09:12:24.313Z",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a>,</p>\n\n<p>I have trained model to follow step-by-step instructions but Unfortunately, the loss always keeps in 11.3~11.3 in the whole training process. It seems nothing to be learned!</p>\n\n<p>Please help me to resolved this issue.</p>",
      "rawMarkdown": "@andrefaraujo,\n\nI have trained model to follow step-by-step instructions but Unfortunately, the loss always keeps in 11.3~11.3 in the whole training process. It seems nothing to be learned!\n\nPlease help me to resolved this issue."
    },
    {
      "id": 915958,
      "postDate": "2020-07-05T08:36:34.693Z",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> <a href=\"/camaskew\">@camaskew</a> What is the expected output type of the model? I've created my own model. and every time I submit it, I get submission error. The output of my model is just a tensor with the shape (1 , 2048), like this one:\n`</p>",
      "rawMarkdown": "@andrefaraujo @camaskew What is the expected output type of the model? I've created my own model. and every time I submit it, I get submission error. The output of my model is just a tensor with the shape (1 , 2048), like this one:\n`",
      "replies": [
        {
          "id": 916491,
          "postDate": "2020-07-05T17:44:35.773Z",
          "content": "<p>See my response in <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163589\">this discussion post</a> for details.</p>",
          "rawMarkdown": "See my response in [this discussion post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163589) for details."
        },
        {
          "id": 916501,
          "postDate": "2020-07-05T17:49:37.240Z",
          "content": "<p>Thank you for your answer. It is clear now 😄 </p>",
          "rawMarkdown": "Thank you for your answer. It is clear now 😄 "
        }
      ]
    },
    {
      "id": 914972,
      "postDate": "2020-07-04T11:52:53.930Z",
      "content": "<p>Hi Araujo, I made my second submission followed by your suggestion. but failed stating that \"Evolution Exception: The Evolution kernel encountered an unexpected error\". I am just using simple ConvNet no DELF with 512 image size. hope this information can be inferred from the model, can u guess where it went wrong. </p>",
      "rawMarkdown": "Hi Araujo, I made my second submission followed by your suggestion. but failed stating that \"Evolution Exception: The Evolution kernel encountered an unexpected error\". I am just using simple ConvNet no DELF with 512 image size. hope this information can be inferred from the model, can u guess where it went wrong. "
    },
    {
      "id": 913904,
      "postDate": "2020-07-03T14:11:11.957Z",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> . Thanks for sharing the thread step by step. But whenever I try executing the build_image_dataset.py, i am getting image_path not defined. on going through the code, the image path s referenced to image paths (*.jpg). Tried changing the input to original competition file path, was getting unable to write to train-0000-00128 error. Can you please advice on how to move forward. </p>",
      "rawMarkdown": "@andrefaraujo . Thanks for sharing the thread step by step. But whenever I try executing the build_image_dataset.py, i am getting image_path not defined. on going through the code, the image path s referenced to image paths (*.jpg). Tried changing the input to original competition file path, was getting unable to write to train-0000-00128 error. Can you please advice on how to move forward. "
    },
    {
      "id": 913872,
      "postDate": "2020-07-03T13:43:05.223Z",
      "content": "<p>thanks for posting this resource!!!.\ni just want to ask that how much time will it take to train a single model through it also is it compatible with colab??.\nand though my question is wage but you can tell me average time consumed under standard condition to train one epoch. thanks <a href=\"/andrefaraujo\">@andrefaraujo</a> </p>",
      "rawMarkdown": "thanks for posting this resource!!!.\ni just want to ask that how much time will it take to train a single model through it also is it compatible with colab??.\nand though my question is wage but you can tell me average time consumed under standard condition to train one epoch. thanks @andrefaraujo "
    }
  ],
  "comments": [
    {
      "id": 941712,
      "author_name": "Bob Wang",
      "author_url": "",
      "post_date": "2020-07-23T11:18:15.193000",
      "content": "<p>Hi, thanks for sharing this. A little bit confused about preparing data for training. When generating tfrecord files, should I leave arg <code>train_clean_csv_path</code>empty if data downloading from kaggle(clean data).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 931214,
      "author_name": "toshi_k",
      "author_url": "",
      "post_date": "2020-07-16T04:15:56.337000",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> It seems that we can't restart training form saved_model.pb (I want to try fine-tuning with the baseline model). Could you share the trained ckpt model which you set to run <code>export_global_model.py</code> ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 932030,
          "author_name": "Andre Araujo",
          "author_url": "",
          "post_date": "2020-07-16T16:27:49.963000",
          "content": "<p>I can't release the checkpoint.</p>\n\n<p>The baseline model was trained with TF-slim, so one idea would be to convert the ResNet101 weights from TF-slim-resnet to Keras-resnet, by looping over the variables and associating the corresponding ones. You then would set the Keras weights directly, using set_weights -- similar to <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/resnet50.py#L421\">this example</a> where we load ImageNet weights. I haven't done it before, but I would expect it to be doable.</p>\n\n<p>Hope this helps!</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 942468,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T18:56:12.173000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 945422,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-07-25T19:48:27.080000",
          "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> loaded model doesn't have reference to its variables. In your example, it is loaded from h5 file.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 919918,
      "author_name": "kaasuu",
      "author_url": "",
      "post_date": "2020-07-08T07:29:57.377000",
      "content": "<p>Hi <a href=\"/andrefaraujo\">@andrefaraujo</a>,</p>\n\n<p>Can we apply your instructions from above and run it in a nb within kaggle? It's not obvious if it we can and my recent attempts point to the answer \"no\".</p>\n\n<p>Are you expecting us to use GCP for downloading data and training?</p>\n\n<p>Thanks</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 947985,
      "author_name": "Fateh Aliyev",
      "author_url": "",
      "post_date": "2020-07-27T15:37:46.893000",
      "content": "<p>Hello, does anyone know how to make sure that every .tar file actually unpacks into a folder?  I run download_dataset.sh and some of the files don't check out with their .md5 checksums and exit.  Does anyone know how to fix this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 942117,
      "author_name": "Suvronil",
      "author_url": "",
      "post_date": "2020-07-23T15:35:45.783000",
      "content": "<p>Hi <a href=\"/andrefaraujo\">@andrefaraujo</a> <a href=\"/bingyicao\">@bingyicao</a> <a href=\"/camaskew\">@camaskew</a>   , I went through the paper related to DELG. I'm hoping for a little clarity on the training procedure , you guys adopted. Do I understand correct, that R101 was trained only till convolution block  5, using Attention loss and reconstruction loss, along with ArcFace loss? Was ArcFace Loss used to train all the rest of the layers or only convolution block 4?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 942585,
          "author_name": "Bingyi Cao",
          "author_url": "",
          "post_date": "2020-07-23T20:58:05.137000",
          "content": "<p>Hey Leo! For the DELG training is end-to-end using training data with image-level label. Yes, you are correct. During the training, we applied attention_loss, reconstruction_loss and ArcFace_loss to optimize the attention network, AutoEncoder and the classifier correspondingly. ArcFace Loss is a loss function modification, so it's not really applied to layers. Let us know if you have further questions!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 934170,
      "author_name": "ouyjb",
      "author_url": "",
      "post_date": "2020-07-18T09:22:21.917000",
      "content": "<p>Thanks for sharing DELF method. Could you show the mAP scores using DELF global model？ I  get 12.7 mAP and I am not sure whether I trained the model correctly.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 920672,
      "author_name": "Jeremy Ma",
      "author_url": "",
      "post_date": "2020-07-08T18:42:25.843000",
      "content": "<p>Hi <a href=\"/andrefaraujo\">@andrefaraujo</a> , \nIs there a timeline for when the other challenges in this workshop going to come out?\nThanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 915994,
      "author_name": "Pardeep Chhokar",
      "author_url": "",
      "post_date": "2020-07-05T09:12:24.313000",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a>,</p>\n\n<p>I have trained model to follow step-by-step instructions but Unfortunately, the loss always keeps in 11.3~11.3 in the whole training process. It seems nothing to be learned!</p>\n\n<p>Please help me to resolved this issue.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 915958,
      "author_name": "Nuno Ferreira",
      "author_url": "",
      "post_date": "2020-07-05T08:36:34.693000",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> <a href=\"/camaskew\">@camaskew</a> What is the expected output type of the model? I've created my own model. and every time I submit it, I get submission error. The output of my model is just a tensor with the shape (1 , 2048), like this one:\n`</p>",
      "votes": 0,
      "replies": [
        {
          "id": 916491,
          "author_name": "Andre Araujo",
          "author_url": "",
          "post_date": "2020-07-05T17:44:35.773000",
          "content": "<p>See my response in <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163589\">this discussion post</a> for details.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916501,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-07-05T17:49:37.240000",
          "content": "<p>Thank you for your answer. It is clear now 😄 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 914972,
      "author_name": "Uday Kumar Gurugubelli",
      "author_url": "",
      "post_date": "2020-07-04T11:52:53.930000",
      "content": "<p>Hi Araujo, I made my second submission followed by your suggestion. but failed stating that \"Evolution Exception: The Evolution kernel encountered an unexpected error\". I am just using simple ConvNet no DELF with 512 image size. hope this information can be inferred from the model, can u guess where it went wrong. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 913904,
      "author_name": "Manoj Prabhakar",
      "author_url": "",
      "post_date": "2020-07-03T14:11:11.957000",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> . Thanks for sharing the thread step by step. But whenever I try executing the build_image_dataset.py, i am getting image_path not defined. on going through the code, the image path s referenced to image paths (*.jpg). Tried changing the input to original competition file path, was getting unable to write to train-0000-00128 error. Can you please advice on how to move forward. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 913872,
      "author_name": "pranshu",
      "author_url": "",
      "post_date": "2020-07-03T13:43:05.223000",
      "content": "<p>thanks for posting this resource!!!.\ni just want to ask that how much time will it take to train a single model through it also is it compatible with colab??.\nand though my question is wage but you can tell me average time consumed under standard condition to train one epoch. thanks <a href=\"/andrefaraujo\">@andrefaraujo</a> </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "911411": "After trying out [your first submission](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163341), now it’s time to train your own landmark retrieval model!\n\nYou could use the [DELF TensorFlow model training codebase](https://github.com/tensorflow/models/tree/master/research/delf/). To train your first ResNet50-based model, you can follow the step-by-step instructions outlined [here](https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training), which cover in detail:\n\n* codebase installation,\n* GLDv2 image downloading,\n* training data preparation, \n* and (finally!) model training.\n\nOnce model training is complete, you can use the [export\\_global\\_model](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py) script to produce a model in the exact format required by the competition. See the command mentioned [here](https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#kaggle-compatible-global-feature-model) which shows how flags should be set.\n\nWith this exported model, you can then follow the instructions in the [first submission post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163341) to upload the model into Kaggle for scoring.\n\nWe hope this provides you a good place to start, with references for model training and concrete examples for model exporting in the right format.\n\nCheers,\n\nAndre\n\nPS1: One small note about the above-mentioned training code example: the default training process is set up to train a model with DELF local features (which is an additional head on top of the ResNet backbone). For the purposes of this year’s landmark retrieval challenge, this local feature head is not required (but it’s also not doing any harm right now). We plan on updating the codebase later with training examples where only the global feature head is present.\n\nPS2: Please note that, while these instructions should work and provide a decent model, they are not currently tuned for optimal performance in this challenge (this codebase is work-in-progress). In particular, right now we observe lower performance using these instructions, compared to when using the above-mentioned first submission (which was trained using a different codebase). In the coming weeks, we will continue to add improvements to the codebase, such as: GeM pooling, ArcFace loss, etc -- which will certainly help improve performance a lot. We also welcome contributions from the community, so feel free to send pull requests!",
    "941712": "Hi, thanks for sharing this. A little bit confused about preparing data for training. When generating tfrecord files, should I leave arg `train_clean_csv_path`empty if data downloading from kaggle(clean data).",
    "931214": "@andrefaraujo It seems that we can't restart training form saved_model.pb (I want to try fine-tuning with the baseline model). Could you share the trained ckpt model which you set to run `export_global_model.py` ?",
    "919918": "Hi @andrefaraujo,\n\nCan we apply your instructions from above and run it in a nb within kaggle? It's not obvious if it we can and my recent attempts point to the answer \"no\".\n\nAre you expecting us to use GCP for downloading data and training?\n\nThanks\n",
    "947985": "Hello, does anyone know how to make sure that every .tar file actually unpacks into a folder?  I run download_dataset.sh and some of the files don't check out with their .md5 checksums and exit.  Does anyone know how to fix this?",
    "942117": "Hi @andrefaraujo @bingyicao @camaskew   , I went through the paper related to DELG. I'm hoping for a little clarity on the training procedure , you guys adopted. Do I understand correct, that R101 was trained only till convolution block  5, using Attention loss and reconstruction loss, along with ArcFace loss? Was ArcFace Loss used to train all the rest of the layers or only convolution block 4?",
    "934170": "Thanks for sharing DELF method. Could you show the mAP scores using DELF global model？ I  get 12.7 mAP and I am not sure whether I trained the model correctly.",
    "920672": "Hi @andrefaraujo , \nIs there a timeline for when the other challenges in this workshop going to come out?\nThanks",
    "915994": "@andrefaraujo,\n\nI have trained model to follow step-by-step instructions but Unfortunately, the loss always keeps in 11.3~11.3 in the whole training process. It seems nothing to be learned!\n\nPlease help me to resolved this issue.",
    "915958": "@andrefaraujo @camaskew What is the expected output type of the model? I've created my own model. and every time I submit it, I get submission error. The output of my model is just a tensor with the shape (1 , 2048), like this one:\n`",
    "914972": "Hi Araujo, I made my second submission followed by your suggestion. but failed stating that \"Evolution Exception: The Evolution kernel encountered an unexpected error\". I am just using simple ConvNet no DELF with 512 image size. hope this information can be inferred from the model, can u guess where it went wrong. ",
    "913904": "@andrefaraujo . Thanks for sharing the thread step by step. But whenever I try executing the build_image_dataset.py, i am getting image_path not defined. on going through the code, the image path s referenced to image paths (*.jpg). Tried changing the input to original competition file path, was getting unable to write to train-0000-00128 error. Can you please advice on how to move forward. ",
    "913872": "thanks for posting this resource!!!.\ni just want to ask that how much time will it take to train a single model through it also is it compatible with colab??.\nand though my question is wage but you can tell me average time consumed under standard condition to train one epoch. thanks @andrefaraujo "
  }
}