{
  "id": 172414,
  "title": "Finetune the Baseline - LB 0.272+ ",
  "url": "/competitions/landmark-retrieval-2020/discussion/172414",
  "author_name": "",
  "post_date": "2020-08-05T01:00:02.481013500Z",
  "votes": 13,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I have not entered this competition but I am considering entering. From discussions, this seems like a difficult competition to beat baseline. Can we beat baseline by loading baseline and training it more (finetuning)?</p>",
  "messages": [
    {
      "id": "958406",
      "postDate": "08/05/2020 01:00:02",
      "content": "<p>I have not entered this competition but I am considering entering. From discussions, this seems like a difficult competition to beat baseline. Can we beat baseline by loading baseline and training it more (finetuning)?</p>",
      "rawMarkdown": "I have not entered this competition but I am considering entering. From discussions, this seems like a difficult competition to beat baseline. Can we beat baseline by loading baseline and training it more (finetuning)?",
      "votes": null
    },
    {
      "id": "958424",
      "postDate": "08/05/2020 01:23:44",
      "content": "<p>I have tried this approach. It seems that the baseline is trained with TF 1.x. You can load it with tf.saved_model.load(), But when you save it again, you get empty variables folder. I have tried a lot, it doesn't work for now.</p>",
      "rawMarkdown": "I have tried this approach. It seems that the baseline is trained with TF 1.x. You can load it with tf.saved_model.load(), But when you save it again, you get empty variables folder. I have tried a lot, it doesn't work for now.",
      "votes": null
    },
    {
      "id": "958931",
      "postDate": "08/05/2020 08:15:14",
      "content": "<p>I'v tried two different approach without any results (for sure there is some mistake in my approach):</p>\n\n<ul>\n<li>Combine resnet101 baseline in my new tf Graph. The problem is that you can combine Garph with tensor but not graph with graph.</li>\n<li>Use resnet101 baseline as feature extractor in preprocessing pipeline. The problem arise when you have to submit because you can't put it inside a tf.function a loaded model beacuse the graph of resnet won't be initialized.</li>\n</ul>",
      "rawMarkdown": "I'v tried two different approach without any results (for sure there is some mistake in my approach):\n\n- Combine resnet101 baseline in my new tf Graph. The problem is that you can combine Garph with tensor but not graph with graph.\n- Use resnet101 baseline as feature extractor in preprocessing pipeline. The problem arise when you have to submit because you can't put it inside a tf.function a loaded model beacuse the graph of resnet won't be initialized.",
      "votes": null
    },
    {
      "id": "959013",
      "postDate": "08/05/2020 09:36:03",
      "content": "<p>Seems like we can't fine tune on savedmodel format which the baseline is in. We would need checkpoint files for that which the organizers won't release.</p>",
      "rawMarkdown": "Seems like we can't fine tune on savedmodel format which the baseline is in. We would need checkpoint files for that which the organizers won't release.",
      "votes": null
    },
    {
      "id": "959327",
      "postDate": "08/05/2020 13:53:18",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I think you have enough compute at your disposal to beat the baseline !</p>",
      "rawMarkdown": "cdeotte I think you have enough compute at your disposal to beat the baseline !",
      "votes": null
    },
    {
      "id": "959362",
      "postDate": "08/05/2020 14:32:35",
      "content": "<p>true, but i feel like there is an easy way to beat baseline. For example, if you have the file on disk which contains 1 million model weights, what would happen if we changed one weight by 0.001 after careful analysis? Could we achieve 0.2710001 and beat baseline?</p>",
      "rawMarkdown": "true, but i feel like there is an easy way to beat baseline. For example, if you have the file on disk which contains 1 million model weights, what would happen if we changed one weight by 0.001 after careful analysis? Could we achieve 0.2710001 and beat baseline?",
      "votes": null
    },
    {
      "id": "959371",
      "postDate": "08/05/2020 14:38:04",
      "content": "<p>I tried this- It doesn't seem possible to read the tensors to change them without the checkpoint file explaining the .data structure. I am new to tf, so perhaps there is a way and I didn't discover it.</p>\n\n<p>It feels to me like the lest worst option is to release the ckpt file, though that is still not super fair to those with scores just above the baseline.</p>",
      "rawMarkdown": "I tried this- It doesn't seem possible to read the tensors to change them without the checkpoint file explaining the .data structure. I am new to tf, so perhaps there is a way and I didn't discover it.\n\nIt feels to me like the lest worst option is to release the ckpt file, though that is still not super fair to those with scores just above the baseline.",
      "votes": null
    },
    {
      "id": "962976",
      "postDate": "08/08/2020 15:20:50",
      "content": "<p>I managed to load weights and re-save them into Keras reimplementation of the model, but it required significant dev effort and discovering a bunch of hacks by reading TF code, so it's possible but definitely non-trivial.</p>\n\n<p>I'm trying the fine-tuning approach now to see whether it is easy-ish to beat baseline that way, but it still requires huge computational resources, to which I have limited access :/</p>",
      "rawMarkdown": "I managed to load weights and re-save them into Keras reimplementation of the model, but it required significant dev effort and discovering a bunch of hacks by reading TF code, so it's possible but definitely non-trivial.\n\nI'm trying the fine-tuning approach now to see whether it is easy-ish to beat baseline that way, but it still requires huge computational resources, to which I have limited access :/",
      "votes": null
    },
    {
      "id": "962978",
      "postDate": "08/08/2020 15:23:01",
      "content": "<p>Great. Note that you can reduce computational resources by freezing layers. Let me know what your results are. (Freezing layers both increases speed and batch size capacity).</p>",
      "rawMarkdown": "Great. Note that you can reduce computational resources by freezing layers. Let me know what your results are. (Freezing layers both increases speed and batch size capacity).",
      "votes": null
    },
    {
      "id": "962988",
      "postDate": "08/08/2020 15:30:21",
      "content": "<p>Well, the biggest problem right there is that if you train on GLDv2 clean dataset, it's about ~80k classes, and when you use ArcFace loss, it requires you to build a dense layer of <code>num_features_in_embedding x num_classes</code> parameters. Since baseline model contains 2048 features, it leads to ArcFace layer containing around 150M parameters, which is 3x as much as feature extractor (ResNet-101) itself.</p>\n\n<p>So fine-tuning with limited resources probably requires to hack around that issue as well.</p>",
      "rawMarkdown": "Well, the biggest problem right there is that if you train on GLDv2 clean dataset, it's about ~80k classes, and when you use ArcFace loss, it requires you to build a dense layer of `num_features_in_embedding x num_classes` parameters. Since baseline model contains 2048 features, it leads to ArcFace layer containing around 150M parameters, which is 3x as much as feature extractor (ResNet-101) itself.\n\nSo fine-tuning with limited resources probably requires to hack around that issue as well.",
      "votes": null
    },
    {
      "id": "963661",
      "postDate": "08/09/2020 07:40:52",
      "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> you can insert another dense layer and have 2048 =&gt; 512 =&gt; 80.000 so the params would be reduced to ~40M</p>",
      "rawMarkdown": "qiubit you can insert another dense layer and have 2048 =&gt; 512 =&gt; 80.000 so the params would be reduced to ~40M",
      "votes": null
    },
    {
      "id": "964588",
      "postDate": "08/10/2020 03:10:15",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": null
    },
    {
      "id": "965266",
      "postDate": "08/10/2020 14:01:57",
      "content": "<p>Is the baseline model trained on the 5M data or the clean (1.5M) data?</p>",
      "rawMarkdown": "Is the baseline model trained on the 5M data or the clean (1.5M) data?",
      "votes": null
    },
    {
      "id": "965499",
      "postDate": "08/10/2020 17:17:54",
      "content": "<p>V2 clean</p>",
      "rawMarkdown": "V2 clean",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 958424,
      "author_name": "wuliaokaola",
      "author_url": "",
      "post_date": "08/05/2020 01:23:44",
      "content": "<p>I have tried this approach. It seems that the baseline is trained with TF 1.x. You can load it with tf.saved_model.load(), But when you save it again, you get empty variables folder. I have tried a lot, it doesn't work for now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 958931,
      "author_name": "stenford23",
      "author_url": "",
      "post_date": "08/05/2020 08:15:14",
      "content": "<p>I'v tried two different approach without any results (for sure there is some mistake in my approach):</p>\n\n<ul>\n<li>Combine resnet101 baseline in my new tf Graph. The problem is that you can combine Garph with tensor but not graph with graph.</li>\n<li>Use resnet101 baseline as feature extractor in preprocessing pipeline. The problem arise when you have to submit because you can't put it inside a tf.function a loaded model beacuse the graph of resnet won't be initialized.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 959013,
      "author_name": "mayukh18",
      "author_url": "",
      "post_date": "08/05/2020 09:36:03",
      "content": "<p>Seems like we can't fine tune on savedmodel format which the baseline is in. We would need checkpoint files for that which the organizers won't release.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 959327,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "08/05/2020 13:53:18",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I think you have enough compute at your disposal to beat the baseline !</p>",
      "votes": null,
      "replies": [
        {
          "id": 959362,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/05/2020 14:32:35",
          "content": "<p>true, but i feel like there is an easy way to beat baseline. For example, if you have the file on disk which contains 1 million model weights, what would happen if we changed one weight by 0.001 after careful analysis? Could we achieve 0.2710001 and beat baseline?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959371,
          "author_name": "robga",
          "author_url": "",
          "post_date": "08/05/2020 14:38:04",
          "content": "<p>I tried this- It doesn't seem possible to read the tensors to change them without the checkpoint file explaining the .data structure. I am new to tf, so perhaps there is a way and I didn't discover it.</p>\n\n<p>It feels to me like the lest worst option is to release the ckpt file, though that is still not super fair to those with scores just above the baseline.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 962976,
      "author_name": "qiubit",
      "author_url": "",
      "post_date": "08/08/2020 15:20:50",
      "content": "<p>I managed to load weights and re-save them into Keras reimplementation of the model, but it required significant dev effort and discovering a bunch of hacks by reading TF code, so it's possible but definitely non-trivial.</p>\n\n<p>I'm trying the fine-tuning approach now to see whether it is easy-ish to beat baseline that way, but it still requires huge computational resources, to which I have limited access :/</p>",
      "votes": null,
      "replies": [
        {
          "id": 962978,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/08/2020 15:23:01",
          "content": "<p>Great. Note that you can reduce computational resources by freezing layers. Let me know what your results are. (Freezing layers both increases speed and batch size capacity).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962988,
          "author_name": "qiubit",
          "author_url": "",
          "post_date": "08/08/2020 15:30:21",
          "content": "<p>Well, the biggest problem right there is that if you train on GLDv2 clean dataset, it's about ~80k classes, and when you use ArcFace loss, it requires you to build a dense layer of <code>num_features_in_embedding x num_classes</code> parameters. Since baseline model contains 2048 features, it leads to ArcFace layer containing around 150M parameters, which is 3x as much as feature extractor (ResNet-101) itself.</p>\n\n<p>So fine-tuning with limited resources probably requires to hack around that issue as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963661,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "08/09/2020 07:40:52",
          "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> you can insert another dense layer and have 2048 =&gt; 512 =&gt; 80.000 so the params would be reduced to ~40M</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964588,
          "author_name": "ranjanp28",
          "author_url": "",
          "post_date": "08/10/2020 03:10:15",
          "content": "<p>yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965266,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "08/10/2020 14:01:57",
          "content": "<p>Is the baseline model trained on the 5M data or the clean (1.5M) data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965499,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "08/10/2020 17:17:54",
          "content": "<p>V2 clean</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "958406": "I have not entered this competition but I am considering entering. From discussions, this seems like a difficult competition to beat baseline. Can we beat baseline by loading baseline and training it more (finetuning)?",
    "958424": "I have tried this approach. It seems that the baseline is trained with TF 1.x. You can load it with tf.saved_model.load(), But when you save it again, you get empty variables folder. I have tried a lot, it doesn't work for now.",
    "958931": "I'v tried two different approach without any results (for sure there is some mistake in my approach):\n\n- Combine resnet101 baseline in my new tf Graph. The problem is that you can combine Garph with tensor but not graph with graph.\n- Use resnet101 baseline as feature extractor in preprocessing pipeline. The problem arise when you have to submit because you can't put it inside a tf.function a loaded model beacuse the graph of resnet won't be initialized.",
    "959013": "Seems like we can't fine tune on savedmodel format which the baseline is in. We would need checkpoint files for that which the organizers won't release.",
    "959327": "cdeotte I think you have enough compute at your disposal to beat the baseline !",
    "959362": "true, but i feel like there is an easy way to beat baseline. For example, if you have the file on disk which contains 1 million model weights, what would happen if we changed one weight by 0.001 after careful analysis? Could we achieve 0.2710001 and beat baseline?",
    "959371": "I tried this- It doesn't seem possible to read the tensors to change them without the checkpoint file explaining the .data structure. I am new to tf, so perhaps there is a way and I didn't discover it.\n\nIt feels to me like the lest worst option is to release the ckpt file, though that is still not super fair to those with scores just above the baseline.",
    "962976": "I managed to load weights and re-save them into Keras reimplementation of the model, but it required significant dev effort and discovering a bunch of hacks by reading TF code, so it's possible but definitely non-trivial.\n\nI'm trying the fine-tuning approach now to see whether it is easy-ish to beat baseline that way, but it still requires huge computational resources, to which I have limited access :/",
    "962978": "Great. Note that you can reduce computational resources by freezing layers. Let me know what your results are. (Freezing layers both increases speed and batch size capacity).",
    "962988": "Well, the biggest problem right there is that if you train on GLDv2 clean dataset, it's about ~80k classes, and when you use ArcFace loss, it requires you to build a dense layer of `num_features_in_embedding x num_classes` parameters. Since baseline model contains 2048 features, it leads to ArcFace layer containing around 150M parameters, which is 3x as much as feature extractor (ResNet-101) itself.\n\nSo fine-tuning with limited resources probably requires to hack around that issue as well.",
    "963661": "qiubit you can insert another dense layer and have 2048 =&gt; 512 =&gt; 80.000 so the params would be reduced to ~40M",
    "964588": "yes",
    "965266": "Is the baseline model trained on the 5M data or the clean (1.5M) data?",
    "965499": "V2 clean"
  },
  "source": "meta"
}