{
  "id": 174283,
  "title": "Why so many teams get map of 0.277? Is it a new baseline? ",
  "url": "/competitions/landmark-retrieval-2020/discussion/174283",
  "author_name": "ouyjb",
  "post_date": "2020-08-13T01:20:34.755000",
  "votes": 9,
  "comment_count": 29,
  "views": 0,
  "content": "<p>The model of 0.271 is from the organizer .<br>\nCould anyone provide some information about the model of 0.277?</p>",
  "messages": [
    {
      "id": 969564,
      "postDate": "2020-08-13T19:13:47.140Z",
      "content": "<p>BTW, I examined the model and it seems to be the Google's baseline feature extractor (it returns VERY similar features as baseline, so I suppose even weights are unchanged), but with the following differences between preprocessing and postprocessing steps:</p>\n<p>Original Google baseline</p>\n<ol>\n<li>x = (input_img - 128.0)/128.0</li>\n<li>height, width = shape(x)</li>\n<li>x1 = resize(x, [(1/np.sqrt(2)) * width, (1/np.sqrt(2)) * height])</li>\n<li>x2 = x</li>\n<li>x3 = resize(x, [np.sqrt(2) * width, np.sqrt(2) * height])</li>\n<li>e1 = embedding_model(x1)</li>\n<li>e2 = embedding_model(x2)</li>\n<li>e3 = embedding_model(x3)</li>\n<li>emb = l2_normalize((e1+e2+e3)/3.0)</li>\n</ol>\n<p>The new 0.277 baseline</p>\n<ol>\n<li>x = (input_img / 128.0) - 1.0</li>\n<li>height, width = shape(x)</li>\n<li>x1 = resize(x, [(1/np.sqrt(2)) * width, (1/np.sqrt(2)) * height])</li>\n<li>x2 = resize(x1, [(1 * width,  1 * height])</li>\n<li>x3 = resize(x2, [np.sqrt(2) * width, np.sqrt(2) * height])</li>\n<li>e1 = embedding_model(x1)</li>\n<li>e2 = embedding_model(x2)</li>\n<li>e3 = embedding_model(x3)</li>\n<li>emb = l2_normalize(e1+e2+e3)</li>\n</ol>\n<p>It's very bizzare that this second pipeline gives higher public leaderboard score, even though the resizing method seems buggy.</p>",
      "rawMarkdown": "BTW, I examined the model and it seems to be the Google's baseline feature extractor (it returns VERY similar features as baseline, so I suppose even weights are unchanged), but with the following differences between preprocessing and postprocessing steps:\n\nOriginal Google baseline\n1. x = (input_img - 128.0)/128.0\n2. height, width = shape(x)\n3. x1 = resize(x, [(1/np.sqrt(2)) * width, (1/np.sqrt(2)) * height])\n4. x2 = x\n5. x3 = resize(x, [np.sqrt(2) * width, np.sqrt(2) * height])\n6. e1 = embedding_model(x1)\n7. e2 = embedding_model(x2)\n8. e3 = embedding_model(x3)\n9. emb = l2_normalize((e1+e2+e3)/3.0)\n\nThe new 0.277 baseline\n1. x = (input_img / 128.0) - 1.0\n2. height, width = shape(x)\n3. x1 = resize(x, [(1/np.sqrt(2)) * width, (1/np.sqrt(2)) * height])\n4. x2 = resize(x1, [(1 * width,  1 * height])\n5. x3 = resize(x2, [np.sqrt(2) * width, np.sqrt(2) * height])\n6. e1 = embedding_model(x1)\n7. e2 = embedding_model(x2)\n8. e3 = embedding_model(x3)\n9. emb = l2_normalize(e1+e2+e3)\n\nIt's very bizzare that this second pipeline gives higher public leaderboard score, even though the resizing method seems buggy.",
      "votes": 11,
      "replies": [
        {
          "id": 969654,
          "postDate": "2020-08-13T20:35:38.293Z",
          "content": "<p>How are you examining these details? Is there any API for this? Not that good at TF and couldn't find any such doc for savedmodels.</p>",
          "rawMarkdown": "How are you examining these details? Is there any API for this? Not that good at TF and couldn't find any such doc for savedmodels."
        },
        {
          "id": 969677,
          "postDate": "2020-08-13T21:02:15.280Z",
          "content": "<p>A combination of using Tensorboard, Python's <code>dir()</code> function (very helpful!), reading TF code and even modifying <code>saved_model.pb</code> by inserting <code>tf.Print</code> operations inside graph.</p>\n<p>As far as I'm concerned, there are no official APIs for doing this, so it's mainly a matter of hacking around TF's hairy internal APIs. From \"official\" things I tried to use tfdbg, but didn't work for me.</p>",
          "rawMarkdown": "A combination of using Tensorboard, Python's `dir()` function (very helpful!), reading TF code and even modifying `saved_model.pb` by inserting `tf.Print` operations inside graph.\n\nAs far as I'm concerned, there are no official APIs for doing this, so it's mainly a matter of hacking around TF's hairy internal APIs. From \"official\" things I tried to use tfdbg, but didn't work for me.",
          "votes": 3
        },
        {
          "id": 969695,
          "postDate": "2020-08-13T21:27:09.323Z",
          "content": "<p>The <code>dir()</code> bit sounds interesting. TF docs are a mess though. </p>\n<p>Initially thought I could load a savedmodel and postprocess but found out savedmodel doesn't really get loaded into memory like a keras/pytorch model, so can't put it inside another savedmodel. Tried so many things, now just resigned.</p>",
          "rawMarkdown": "The `dir()` bit sounds interesting. TF docs are a mess though. \n\nInitially thought I could load a savedmodel and postprocess but found out savedmodel doesn't really get loaded into memory like a keras/pytorch model, so can't put it inside another savedmodel. Tried so many things, now just resigned."
        },
        {
          "id": 969829,
          "postDate": "2020-08-14T01:27:09.793Z",
          "content": "<p>You are doing me a great service, and I'm very grateful to you!</p>",
          "rawMarkdown": "You are doing me a great service, and I'm very grateful to you!"
        },
        {
          "id": 970015,
          "postDate": "2020-08-14T06:06:11.723Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 970050,
          "postDate": "2020-08-14T06:43:52.330Z",
          "content": "<p>It's very interesting - how did you get the conclusion about VGG?</p>\n<p>Here's a notebook which shows the experiment after which I concluded both embedding models are the same - <a href=\"https://www.kaggle.com/qiubit/comparison-0-271-vs-0-277\" target=\"_blank\">https://www.kaggle.com/qiubit/comparison-0-271-vs-0-277</a></p>",
          "rawMarkdown": "It's very interesting - how did you get the conclusion about VGG?\n\nHere's a notebook which shows the experiment after which I concluded both embedding models are the same - https://www.kaggle.com/qiubit/comparison-0-271-vs-0-277",
          "votes": 1,
          "replies": [
            {
              "id": 970053,
              "postDate": "2020-08-14T06:46:34.567Z",
              "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> Does that mean both 0.277 and 0.271 should have the same private LB score, or rather should have the same Public LB score as well?</p>",
              "rawMarkdown": "@qiubit Does that mean both 0.277 and 0.271 should have the same private LB score, or rather should have the same Public LB score as well?"
            },
            {
              "id": 970054,
              "postDate": "2020-08-14T06:47:56.383Z",
              "content": "<p>No - because of different preprocessing and postprocessing steps (saved model contains entire preprocessing and postprocessing pipeline, in the notebook above I only show networks that, I think, are used for creating feature embeddings).</p>",
              "rawMarkdown": "No - because of different preprocessing and postprocessing steps (saved model contains entire preprocessing and postprocessing pipeline, in the notebook above I only show networks that, I think, are used for creating feature embeddings).",
              "votes": 1
            },
            {
              "id": 970057,
              "postDate": "2020-08-14T06:49:19.493Z",
              "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> Got it, but from what you show their model structure should be the same. But I'm not an expert in this yet so i'm not sure</p>",
              "rawMarkdown": "@qiubit Got it, but from what you show their model structure should be the same. But I'm not an expert in this yet so i'm not sure"
            },
            {
              "id": 970192,
              "postDate": "2020-08-14T09:18:30.097Z",
              "content": "<p>Sorry, nevermind my previous (deleted) comment, I loaded the wrong model.</p>",
              "rawMarkdown": "Sorry, nevermind my previous (deleted) comment, I loaded the wrong model."
            }
          ]
        },
        {
          "id": 970076,
          "postDate": "2020-08-14T07:18:42.223Z",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> . Can you also share your script to examine the model ? </p>",
          "rawMarkdown": "Thanks for sharing @qiubit . Can you also share your script to examine the model ? "
        },
        {
          "id": 970093,
          "postDate": "2020-08-14T07:34:02.797Z",
          "content": "<p>I don't have a single script, just a bunch of experiments scattered all over the place, so unfortunately nothing shareable straight away :&lt;</p>",
          "rawMarkdown": "I don't have a single script, just a bunch of experiments scattered all over the place, so unfortunately nothing shareable straight away :<",
          "votes": 1
        },
        {
          "id": 970208,
          "postDate": "2020-08-14T09:27:04.320Z",
          "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> No worries</p>",
          "rawMarkdown": "@qiubit No worries"
        },
        {
          "id": 970272,
          "postDate": "2020-08-14T10:04:32.613Z",
          "content": "<p>I wonder what is the essential difference between these two models.<br>\nAt least, 'emb = l2_normalize((e1+e2+e3)/3.0)' and 'l2_normalize(e1+e2+e3)' output the same vectors by normalization, don't them ?</p>",
          "rawMarkdown": "I wonder what is the essential difference between these two models.\nAt least, 'emb = l2_normalize((e1+e2+e3)/3.0)' and 'l2_normalize(e1+e2+e3)' output the same vectors by normalization, don't them ?",
          "votes": 1,
          "replies": [
            {
              "id": 970279,
              "postDate": "2020-08-14T10:08:45.820Z",
              "content": "<p>I think the preprocessing step is the main difference [different image gets resized in model that gives 0.277 score].</p>",
              "rawMarkdown": "I think the preprocessing step is the main difference [different image gets resized in model that gives 0.277 score]."
            },
            {
              "id": 971000,
              "postDate": "2020-08-15T04:51:18.737Z",
              "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> The preprocessing step is quite similar in your pseudocode, I would be surprised if that makes such large difference. There is a difference between <code>l2_normalize((e1+e2+e3)/3.0)</code> and <code>l2_normalize(e1+e2+e3)</code> in some situations (at least I accidentally caught one in one of my models). It has to do with <a href=\"https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html#:~:text=Floating%2Dpoint%20representations%20have%20a,approximately%201.10011001100110011001101%20%C3%97%202%2D4.\" target=\"_blank\">floating points rounding error</a>. In my case I caught a difference of <code>0.0002</code> accumulated over all coordinates.</p>",
              "rawMarkdown": "@qiubit The preprocessing step is quite similar in your pseudocode, I would be surprised if that makes such large difference. There is a difference between `l2_normalize((e1+e2+e3)/3.0)` and `l2_normalize(e1+e2+e3)` in some situations (at least I accidentally caught one in one of my models). It has to do with [floating points rounding error](https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html#:~:text=Floating%2Dpoint%20representations%20have%20a,approximately%201.10011001100110011001101%20%C3%97%202%2D4.). In my case I caught a difference of `0.0002` accumulated over all coordinates."
            }
          ]
        },
        {
          "id": 970781,
          "postDate": "2020-08-14T19:34:57.237Z",
          "content": "<p>Why are the <code>width</code> and <code>height</code> reversed in <code>resize</code> function?</p>",
          "rawMarkdown": "Why are the `width` and `height` reversed in `resize` function?",
          "votes": 1
        },
        {
          "id": 970796,
          "postDate": "2020-08-14T19:59:42.170Z",
          "content": "<p>It's pseudocode, the details are not 1-1 with tensorflow. It's correct in the proper model</p>",
          "rawMarkdown": "It's pseudocode, the details are not 1-1 with tensorflow. It's correct in the proper model",
          "votes": 1
        },
        {
          "id": 970876,
          "postDate": "2020-08-14T23:33:39.417Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 971157,
          "postDate": "2020-08-15T08:18:56.753Z",
          "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> Wondering if the 0.277 has decimal places in the LB that is unseen - not sure if each sub of 0.277 has different decimal places behind…</p>",
          "rawMarkdown": "@qiubit Wondering if the 0.277 has decimal places in the LB that is unseen - not sure if each sub of 0.277 has different decimal places behind..."
        }
      ]
    },
    {
      "id": 968399,
      "postDate": "2020-08-13T01:20:34.757Z",
      "content": "<p>The model of 0.271 is from the organizer .<br>\nCould anyone provide some information about the model of 0.277?</p>",
      "rawMarkdown": "The model of 0.271 is from the organizer .\nCould anyone provide some information about the model of 0.277?\n",
      "votes": 9
    },
    {
      "id": 974801,
      "postDate": "2020-08-18T03:59:11.320Z",
      "content": "<p>An observation - I just scaled up the input size for 0.277 baseline by a factor sqrt(2) and it improved my Public LB score to 0.281 and worked well on Private LB too 😄.</p>\n<p><a href=\"https://www.kaggle.com/prateekagnihotri/change-image-size-to-get-0-281-lb\" target=\"_blank\">This</a> is the link where you can find this notebook.</p>",
      "rawMarkdown": "An observation - I just scaled up the input size for 0.277 baseline by a factor sqrt(2) and it improved my Public LB score to 0.281 and worked well on Private LB too 😄.\n\n[This](https://www.kaggle.com/prateekagnihotri/change-image-size-to-get-0-281-lb) is the link where you can find this notebook.",
      "votes": 5
    },
    {
      "id": 968846,
      "postDate": "2020-08-13T09:23:44.533Z",
      "content": "<p>Actually, that notebook was published earlier, then it was removed. I just made it published again to avoid private sharing.</p>",
      "rawMarkdown": "Actually, that notebook was published earlier, then it was removed. I just made it published again to avoid private sharing.",
      "votes": 4,
      "replies": [
        {
          "id": 968862,
          "postDate": "2020-08-13T09:39:35.277Z",
          "content": "<p><a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> What does Kaggle do when there are so many ties in score though… Will medals still be given in such a tie?</p>",
          "rawMarkdown": "@nvnnghia What does Kaggle do when there are so many ties in score though... Will medals still be given in such a tie?",
          "replies": [
            {
              "id": 968868,
              "postDate": "2020-08-13T09:48:04.330Z",
              "content": "<p>who submit first will have a higher rank in case of equal scores. </p>",
              "rawMarkdown": "who submit first will have a higher rank in case of equal scores. ",
              "votes": 1
            },
            {
              "id": 968874,
              "postDate": "2020-08-13T09:58:55.050Z",
              "content": "<p><a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> I see, thank you for the reply!</p>",
              "rawMarkdown": "@nvnnghia I see, thank you for the reply!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 968436,
      "postDate": "2020-08-13T02:36:28.427Z",
      "content": "<p>You can find that notebook by going to notebooks and sorting by \"best score\". Its this one: <a href=\"https://www.kaggle.com/nvnnghia/main-0806\" target=\"_blank\">https://www.kaggle.com/nvnnghia/main-0806</a></p>",
      "rawMarkdown": "You can find that notebook by going to notebooks and sorting by \"best score\". Its this one: https://www.kaggle.com/nvnnghia/main-0806",
      "votes": 1,
      "replies": [
        {
          "id": 968665,
          "postDate": "2020-08-13T06:58:47.387Z",
          "content": "<p>Do you have any information about this model ?</p>",
          "rawMarkdown": "Do you have any information about this model ?"
        }
      ]
    },
    {
      "id": 968663,
      "postDate": "2020-08-13T06:58:31.653Z",
      "content": "<p>Do you have any information about this model ?</p>",
      "rawMarkdown": "Do you have any information about this model ?"
    }
  ],
  "comments": [
    {
      "id": 969564,
      "author_name": "qiubit",
      "author_url": "",
      "post_date": "2020-08-13T19:13:47.140000",
      "content": "<p>BTW, I examined the model and it seems to be the Google's baseline feature extractor (it returns VERY similar features as baseline, so I suppose even weights are unchanged), but with the following differences between preprocessing and postprocessing steps:</p>\n<p>Original Google baseline</p>\n<ol>\n<li>x = (input_img - 128.0)/128.0</li>\n<li>height, width = shape(x)</li>\n<li>x1 = resize(x, [(1/np.sqrt(2)) * width, (1/np.sqrt(2)) * height])</li>\n<li>x2 = x</li>\n<li>x3 = resize(x, [np.sqrt(2) * width, np.sqrt(2) * height])</li>\n<li>e1 = embedding_model(x1)</li>\n<li>e2 = embedding_model(x2)</li>\n<li>e3 = embedding_model(x3)</li>\n<li>emb = l2_normalize((e1+e2+e3)/3.0)</li>\n</ol>\n<p>The new 0.277 baseline</p>\n<ol>\n<li>x = (input_img / 128.0) - 1.0</li>\n<li>height, width = shape(x)</li>\n<li>x1 = resize(x, [(1/np.sqrt(2)) * width, (1/np.sqrt(2)) * height])</li>\n<li>x2 = resize(x1, [(1 * width,  1 * height])</li>\n<li>x3 = resize(x2, [np.sqrt(2) * width, np.sqrt(2) * height])</li>\n<li>e1 = embedding_model(x1)</li>\n<li>e2 = embedding_model(x2)</li>\n<li>e3 = embedding_model(x3)</li>\n<li>emb = l2_normalize(e1+e2+e3)</li>\n</ol>\n<p>It's very bizzare that this second pipeline gives higher public leaderboard score, even though the resizing method seems buggy.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 969654,
          "author_name": "Mayukh Bhattacharyya",
          "author_url": "",
          "post_date": "2020-08-13T20:35:38.293000",
          "content": "<p>How are you examining these details? Is there any API for this? Not that good at TF and couldn't find any such doc for savedmodels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 969677,
          "author_name": "qiubit",
          "author_url": "",
          "post_date": "2020-08-13T21:02:15.280000",
          "content": "<p>A combination of using Tensorboard, Python's <code>dir()</code> function (very helpful!), reading TF code and even modifying <code>saved_model.pb</code> by inserting <code>tf.Print</code> operations inside graph.</p>\n<p>As far as I'm concerned, there are no official APIs for doing this, so it's mainly a matter of hacking around TF's hairy internal APIs. From \"official\" things I tried to use tfdbg, but didn't work for me.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 969695,
          "author_name": "Mayukh Bhattacharyya",
          "author_url": "",
          "post_date": "2020-08-13T21:27:09.323000",
          "content": "<p>The <code>dir()</code> bit sounds interesting. TF docs are a mess though. </p>\n<p>Initially thought I could load a savedmodel and postprocess but found out savedmodel doesn't really get loaded into memory like a keras/pytorch model, so can't put it inside another savedmodel. Tried so many things, now just resigned.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 969829,
          "author_name": "ouyjb",
          "author_url": "",
          "post_date": "2020-08-14T01:27:09.793000",
          "content": "<p>You are doing me a great service, and I'm very grateful to you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 970015,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-14T06:06:11.723000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 970050,
          "author_name": "qiubit",
          "author_url": "",
          "post_date": "2020-08-14T06:43:52.330000",
          "content": "<p>It's very interesting - how did you get the conclusion about VGG?</p>\n<p>Here's a notebook which shows the experiment after which I concluded both embedding models are the same - <a href=\"https://www.kaggle.com/qiubit/comparison-0-271-vs-0-277\" target=\"_blank\">https://www.kaggle.com/qiubit/comparison-0-271-vs-0-277</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 970053,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-14T06:46:34.567000",
              "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> Does that mean both 0.277 and 0.271 should have the same private LB score, or rather should have the same Public LB score as well?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 970054,
              "author_name": "qiubit",
              "author_url": "",
              "post_date": "2020-08-14T06:47:56.383000",
              "content": "<p>No - because of different preprocessing and postprocessing steps (saved model contains entire preprocessing and postprocessing pipeline, in the notebook above I only show networks that, I think, are used for creating feature embeddings).</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 970057,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-14T06:49:19.493000",
              "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> Got it, but from what you show their model structure should be the same. But I'm not an expert in this yet so i'm not sure</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 970192,
              "author_name": "Chan Kha Vu",
              "author_url": "",
              "post_date": "2020-08-14T09:18:30.097000",
              "content": "<p>Sorry, nevermind my previous (deleted) comment, I loaded the wrong model.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 970076,
          "author_name": "nvnn",
          "author_url": "",
          "post_date": "2020-08-14T07:18:42.223000",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> . Can you also share your script to examine the model ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 970093,
          "author_name": "qiubit",
          "author_url": "",
          "post_date": "2020-08-14T07:34:02.797000",
          "content": "<p>I don't have a single script, just a bunch of experiments scattered all over the place, so unfortunately nothing shareable straight away :&lt;</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 970208,
          "author_name": "nvnn",
          "author_url": "",
          "post_date": "2020-08-14T09:27:04.320000",
          "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> No worries</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 970272,
          "author_name": "toshi_k",
          "author_url": "",
          "post_date": "2020-08-14T10:04:32.613000",
          "content": "<p>I wonder what is the essential difference between these two models.<br>\nAt least, 'emb = l2_normalize((e1+e2+e3)/3.0)' and 'l2_normalize(e1+e2+e3)' output the same vectors by normalization, don't them ?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 970279,
              "author_name": "qiubit",
              "author_url": "",
              "post_date": "2020-08-14T10:08:45.820000",
              "content": "<p>I think the preprocessing step is the main difference [different image gets resized in model that gives 0.277 score].</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 971000,
              "author_name": "Chan Kha Vu",
              "author_url": "",
              "post_date": "2020-08-15T04:51:18.737000",
              "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> The preprocessing step is quite similar in your pseudocode, I would be surprised if that makes such large difference. There is a difference between <code>l2_normalize((e1+e2+e3)/3.0)</code> and <code>l2_normalize(e1+e2+e3)</code> in some situations (at least I accidentally caught one in one of my models). It has to do with <a href=\"https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html#:~:text=Floating%2Dpoint%20representations%20have%20a,approximately%201.10011001100110011001101%20%C3%97%202%2D4.\" target=\"_blank\">floating points rounding error</a>. In my case I caught a difference of <code>0.0002</code> accumulated over all coordinates.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 970781,
          "author_name": "Xiren Zhou",
          "author_url": "",
          "post_date": "2020-08-14T19:34:57.237000",
          "content": "<p>Why are the <code>width</code> and <code>height</code> reversed in <code>resize</code> function?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 970796,
          "author_name": "qiubit",
          "author_url": "",
          "post_date": "2020-08-14T19:59:42.170000",
          "content": "<p>It's pseudocode, the details are not 1-1 with tensorflow. It's correct in the proper model</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 970876,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-14T23:33:39.417000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 971157,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2020-08-15T08:18:56.753000",
          "content": "<p><a href=\"https://www.kaggle.com/qiubit\" target=\"_blank\">@qiubit</a> Wondering if the 0.277 has decimal places in the LB that is unseen - not sure if each sub of 0.277 has different decimal places behind…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974801,
      "author_name": "Prateek",
      "author_url": "",
      "post_date": "2020-08-18T03:59:11.320000",
      "content": "<p>An observation - I just scaled up the input size for 0.277 baseline by a factor sqrt(2) and it improved my Public LB score to 0.281 and worked well on Private LB too 😄.</p>\n<p><a href=\"https://www.kaggle.com/prateekagnihotri/change-image-size-to-get-0-281-lb\" target=\"_blank\">This</a> is the link where you can find this notebook.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 968846,
      "author_name": "nvnn",
      "author_url": "",
      "post_date": "2020-08-13T09:23:44.533000",
      "content": "<p>Actually, that notebook was published earlier, then it was removed. I just made it published again to avoid private sharing.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 968862,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2020-08-13T09:39:35.277000",
          "content": "<p><a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> What does Kaggle do when there are so many ties in score though… Will medals still be given in such a tie?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 968868,
              "author_name": "nvnn",
              "author_url": "",
              "post_date": "2020-08-13T09:48:04.330000",
              "content": "<p>who submit first will have a higher rank in case of equal scores. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 968874,
              "author_name": "gao-hongnan",
              "author_url": "",
              "post_date": "2020-08-13T09:58:55.050000",
              "content": "<p><a href=\"https://www.kaggle.com/nvnnghia\" target=\"_blank\">@nvnnghia</a> I see, thank you for the reply!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 968436,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2020-08-13T02:36:28.427000",
      "content": "<p>You can find that notebook by going to notebooks and sorting by \"best score\". Its this one: <a href=\"https://www.kaggle.com/nvnnghia/main-0806\" target=\"_blank\">https://www.kaggle.com/nvnnghia/main-0806</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 968665,
          "author_name": "ouyjb",
          "author_url": "",
          "post_date": "2020-08-13T06:58:47.387000",
          "content": "<p>Do you have any information about this model ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 968663,
      "author_name": "ouyjb",
      "author_url": "",
      "post_date": "2020-08-13T06:58:31.653000",
      "content": "<p>Do you have any information about this model ?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "969564": "BTW, I examined the model and it seems to be the Google's baseline feature extractor (it returns VERY similar features as baseline, so I suppose even weights are unchanged), but with the following differences between preprocessing and postprocessing steps:\n\nOriginal Google baseline\n1. x = (input_img - 128.0)/128.0\n2. height, width = shape(x)\n3. x1 = resize(x, [(1/np.sqrt(2)) * width, (1/np.sqrt(2)) * height])\n4. x2 = x\n5. x3 = resize(x, [np.sqrt(2) * width, np.sqrt(2) * height])\n6. e1 = embedding_model(x1)\n7. e2 = embedding_model(x2)\n8. e3 = embedding_model(x3)\n9. emb = l2_normalize((e1+e2+e3)/3.0)\n\nThe new 0.277 baseline\n1. x = (input_img / 128.0) - 1.0\n2. height, width = shape(x)\n3. x1 = resize(x, [(1/np.sqrt(2)) * width, (1/np.sqrt(2)) * height])\n4. x2 = resize(x1, [(1 * width,  1 * height])\n5. x3 = resize(x2, [np.sqrt(2) * width, np.sqrt(2) * height])\n6. e1 = embedding_model(x1)\n7. e2 = embedding_model(x2)\n8. e3 = embedding_model(x3)\n9. emb = l2_normalize(e1+e2+e3)\n\nIt's very bizzare that this second pipeline gives higher public leaderboard score, even though the resizing method seems buggy.",
    "968399": "The model of 0.271 is from the organizer .\nCould anyone provide some information about the model of 0.277?\n",
    "974801": "An observation - I just scaled up the input size for 0.277 baseline by a factor sqrt(2) and it improved my Public LB score to 0.281 and worked well on Private LB too 😄.\n\n[This](https://www.kaggle.com/prateekagnihotri/change-image-size-to-get-0-281-lb) is the link where you can find this notebook.",
    "968846": "Actually, that notebook was published earlier, then it was removed. I just made it published again to avoid private sharing.",
    "968436": "You can find that notebook by going to notebooks and sorting by \"best score\". Its this one: https://www.kaggle.com/nvnnghia/main-0806",
    "968663": "Do you have any information about this model ?"
  }
}