{
  "id": 172576,
  "title": "Host Baseline Example",
  "url": "/competitions/landmark-recognition-2020/discussion/172576",
  "author_name": "Cam Askew",
  "post_date": "2020-08-05T16:09:52.815000",
  "votes": 50,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Howdy, folks!</p>\n\n<p>Since this is a code competition, you'll be submitting a Notebook instead of a csv, as in previous years. This means you can use techniques like ensembles and local feature re-ranking and whatever implementation you like, so long as your kernel runs within the hardware constraints of the Notebook worker and finishes in 12h.</p>\n\n<p>The format <em>can</em> make that first submission a little bit more difficult though, so I put together an <a href=\"https://www.kaggle.com/camaskew/host-baseline-example\">example Notebook</a> and shared it publicly to get you started. There's a more detailed description in the Notebook itself, but in short, it uses TensorFlow to load pre-trained <a href=\"https://arxiv.org/abs/2001.05027\">DELG</a> models from an attached dataset for feature extraction. It makes predictions using retrieval by global feature similarity and re-ranking by geometric verification.</p>\n\n<p>To stay below the time limit, I only extracted local features for 3 scales, and only re-ranked the top <code>NUM_TO_RERANK</code> similarity-ranked training images, using <a href=\"https://pypi.org/project/pydegensac/\">pydegensac</a> for speedy geometric-verification. With <code>NUM_TO_RERANK = 10</code>, runtime is just under 10h. As a heads-up: subsequent re-runs may produce slightly different results, since there's not an easy way to specify a seed for RANSAC.\nThe code also has a few lines that simplified my submission workflow, and I hope they can help you, too! The dataset for the competition is split into public and private subsets. When your code runs in a Draft Session or a Commit, you'll be given the public images to test and train with. Submissions are run against the private set, though, and since you start a submission by clicking \"Submit\" after a Commit outputs <code>submission.csv</code>, the example quickly creates a submission by copying <code>sample_submission.csv</code> from the competition dataset. This way, you don't have to run your code on both the private and public datasets to see your actual score.</p>\n\n<p>I hope the example ends up being helpful, and I can't wait to see your improvements on it!</p>\n\n<p>Happy coding,\nCam</p>",
  "messages": [
    {
      "id": 959487,
      "postDate": "2020-08-05T16:09:52.817Z",
      "content": "<p>Howdy, folks!</p>\n\n<p>Since this is a code competition, you'll be submitting a Notebook instead of a csv, as in previous years. This means you can use techniques like ensembles and local feature re-ranking and whatever implementation you like, so long as your kernel runs within the hardware constraints of the Notebook worker and finishes in 12h.</p>\n\n<p>The format <em>can</em> make that first submission a little bit more difficult though, so I put together an <a href=\"https://www.kaggle.com/camaskew/host-baseline-example\">example Notebook</a> and shared it publicly to get you started. There's a more detailed description in the Notebook itself, but in short, it uses TensorFlow to load pre-trained <a href=\"https://arxiv.org/abs/2001.05027\">DELG</a> models from an attached dataset for feature extraction. It makes predictions using retrieval by global feature similarity and re-ranking by geometric verification.</p>\n\n<p>To stay below the time limit, I only extracted local features for 3 scales, and only re-ranked the top <code>NUM_TO_RERANK</code> similarity-ranked training images, using <a href=\"https://pypi.org/project/pydegensac/\">pydegensac</a> for speedy geometric-verification. With <code>NUM_TO_RERANK = 10</code>, runtime is just under 10h. As a heads-up: subsequent re-runs may produce slightly different results, since there's not an easy way to specify a seed for RANSAC.\nThe code also has a few lines that simplified my submission workflow, and I hope they can help you, too! The dataset for the competition is split into public and private subsets. When your code runs in a Draft Session or a Commit, you'll be given the public images to test and train with. Submissions are run against the private set, though, and since you start a submission by clicking \"Submit\" after a Commit outputs <code>submission.csv</code>, the example quickly creates a submission by copying <code>sample_submission.csv</code> from the competition dataset. This way, you don't have to run your code on both the private and public datasets to see your actual score.</p>\n\n<p>I hope the example ends up being helpful, and I can't wait to see your improvements on it!</p>\n\n<p>Happy coding,\nCam</p>",
      "rawMarkdown": "Howdy, folks!\n\nSince this is a code competition, you'll be submitting a Notebook instead of a csv, as in previous years. This means you can use techniques like ensembles and local feature re-ranking and whatever implementation you like, so long as your kernel runs within the hardware constraints of the Notebook worker and finishes in 12h.\n\nThe format *can* make that first submission a little bit more difficult though, so I put together an [example Notebook](https://www.kaggle.com/camaskew/host-baseline-example) and shared it publicly to get you started. There's a more detailed description in the Notebook itself, but in short, it uses TensorFlow to load pre-trained [DELG](https://arxiv.org/abs/2001.05027) models from an attached dataset for feature extraction. It makes predictions using retrieval by global feature similarity and re-ranking by geometric verification.\n\nTo stay below the time limit, I only extracted local features for 3 scales, and only re-ranked the top `NUM_TO_RERANK` similarity-ranked training images, using [pydegensac](https://pypi.org/project/pydegensac/) for speedy geometric-verification. With `NUM_TO_RERANK = 10`, runtime is just under 10h. As a heads-up: subsequent re-runs may produce slightly different results, since there's not an easy way to specify a seed for RANSAC.\nThe code also has a few lines that simplified my submission workflow, and I hope they can help you, too! The dataset for the competition is split into public and private subsets. When your code runs in a Draft Session or a Commit, you'll be given the public images to test and train with. Submissions are run against the private set, though, and since you start a submission by clicking \"Submit\" after a Commit outputs `submission.csv`, the example quickly creates a submission by copying `sample_submission.csv` from the competition dataset. This way, you don't have to run your code on both the private and public datasets to see your actual score.\n\nI hope the example ends up being helpful, and I can't wait to see your improvements on it!\n\nHappy coding,\nCam\n",
      "votes": 50
    },
    {
      "id": 962022,
      "postDate": "2020-08-07T18:09:55.570Z",
      "content": "<p>it threw an exception when i submit, any idea why?</p>",
      "rawMarkdown": "it threw an exception when i submit, any idea why?",
      "votes": 4,
      "replies": [
        {
          "id": 969364,
          "postDate": "2020-08-13T16:26:15.717Z",
          "content": "<p>Hey, same happened to me… did you find out the reason?</p>",
          "rawMarkdown": "Hey, same happened to me... did you find out the reason?",
          "votes": 1
        },
        {
          "id": 976701,
          "postDate": "2020-08-19T03:38:58.213Z",
          "content": "<p>Same - this base kernel just give an error…</p>",
          "rawMarkdown": "Same - this base kernel just give an error...",
          "votes": 1
        },
        {
          "id": 977249,
          "postDate": "2020-08-19T11:25:45.523Z",
          "content": "<p>I face the same issue here</p>",
          "rawMarkdown": "I face the same issue here"
        },
        {
          "id": 999308,
          "postDate": "2020-09-05T14:54:36.807Z",
          "content": "<p>I am able to submit successfully several days ago using this kernel (version 9) <a href=\"https://www.kaggle.com/paulorzp/baseline-landmark-recognition-lb-0-48?scriptVersionId=41322107\" target=\"_blank\">https://www.kaggle.com/paulorzp/baseline-landmark-recognition-lb-0-48?scriptVersionId=41322107</a> I think it is only differ from the baseline in just few parameter values. I think also submission failure is due to time exceeding of the kernel.</p>",
          "rawMarkdown": "I am able to submit successfully several days ago using this kernel (version 9) https://www.kaggle.com/paulorzp/baseline-landmark-recognition-lb-0-48?scriptVersionId=41322107 I think it is only differ from the baseline in just few parameter values. I think also submission failure is due to time exceeding of the kernel."
        }
      ]
    },
    {
      "id": 1011206,
      "postDate": "2020-09-15T10:03:53.733Z",
      "content": "<p>Thanks for providing the notebook. It was really very helpful!</p>",
      "rawMarkdown": "Thanks for providing the notebook. It was really very helpful!",
      "votes": 1
    },
    {
      "id": 979524,
      "postDate": "2020-08-20T23:10:19.573Z",
      "content": "<p>I run the base line notebook but I got a submission score error, I see also some comments reporting errors while running the base line notebook.  Did anyone able to run the base line successfully ?</p>",
      "rawMarkdown": "I run the base line notebook but I got a submission score error, I see also some comments reporting errors while running the base line notebook.  Did anyone able to run the base line successfully ?",
      "votes": 2
    },
    {
      "id": 1003358,
      "postDate": "2020-09-08T21:46:29.037Z",
      "content": "<p>Thank you for the example. <br>\nIn the public train data set, every landmark presented at least twice. Is it correct also for private train set?</p>",
      "rawMarkdown": "Thank you for the example. \nIn the public train data set, every landmark presented at least twice. Is it correct also for private train set?"
    },
    {
      "id": 993939,
      "postDate": "2020-09-01T08:43:44.357Z",
      "content": "<p>Please help me with the following error. </p>\n<p>OSError: SavedModel file does not exist at: ../input/delg-saved-models/local_and_global/{saved_model.pbtxt|saved_model.pb}</p>",
      "rawMarkdown": "Please help me with the following error. \n\nOSError: SavedModel file does not exist at: ../input/delg-saved-models/local_and_global/{saved_model.pbtxt|saved_model.pb}",
      "replies": [
        {
          "id": 996895,
          "postDate": "2020-09-03T16:18:25.217Z",
          "content": "<p>Try correcting the filename with the file format extension.<br>\nThanks</p>",
          "rawMarkdown": "Try correcting the filename with the file format extension.\nThanks"
        }
      ]
    },
    {
      "id": 965316,
      "postDate": "2020-08-10T14:42:55.593Z",
      "content": "<p>I was wondering if the baseline model given in your example an exact copy of the DELG pre-trained on the Google-Landmarks dataset v1 or the RN101-ArcFace pre-trained on Google-Landmarks dataset v2 (clean)?</p>",
      "rawMarkdown": "I was wondering if the baseline model given in your example an exact copy of the DELG pre-trained on the Google-Landmarks dataset v1 or the RN101-ArcFace pre-trained on Google-Landmarks dataset v2 (clean)?",
      "replies": [
        {
          "id": 966568,
          "postDate": "2020-08-11T14:28:48.330Z",
          "content": "<p>Yep, the weights are identical to the publicly available pre-trained model.</p>",
          "rawMarkdown": "Yep, the weights are identical to the publicly available pre-trained model."
        }
      ]
    },
    {
      "id": 961054,
      "postDate": "2020-08-06T21:55:10.480Z",
      "content": "<p>Will you be releasing the training code?</p>",
      "rawMarkdown": "Will you be releasing the training code?",
      "replies": [
        {
          "id": 961058,
          "postDate": "2020-08-06T22:04:14.900Z",
          "content": "<p>Take a look at <a href=\"https://github.com/tensorflow/models/tree/master/research/delf\">the github repo</a>!</p>",
          "rawMarkdown": "Take a look at [the github repo](https://github.com/tensorflow/models/tree/master/research/delf)!",
          "votes": 1
        }
      ]
    },
    {
      "id": 960866,
      "postDate": "2020-08-06T18:45:51.793Z",
      "content": "<p>Thank you for the baseline example! I would like to ask if the ~10h runtime is achieved on gpu?</p>",
      "rawMarkdown": "Thank you for the baseline example! I would like to ask if the ~10h runtime is achieved on gpu?",
      "replies": [
        {
          "id": 960963,
          "postDate": "2020-08-06T20:24:00.357Z",
          "content": "<p>yup using GPU</p>",
          "rawMarkdown": "yup using GPU"
        }
      ]
    },
    {
      "id": 960864,
      "postDate": "2020-08-06T18:42:37.173Z",
      "content": "<p>Thanks for providing this notebook! Do you have any resources to get started with image retrieval? 👀 </p>",
      "rawMarkdown": "Thanks for providing this notebook! Do you have any resources to get started with image retrieval? 👀 "
    },
    {
      "id": 1009947,
      "postDate": "2020-09-14T11:48:22.867Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 999185,
      "postDate": "2020-09-05T12:58:27.160Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 998081,
      "postDate": "2020-09-04T13:30:48.517Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 976323,
      "postDate": "2020-08-18T19:36:27.083Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 963859,
      "postDate": "2020-08-09T11:17:51.240Z",
      "content": "<p>Thanks for promoting pydegensac!</p>",
      "rawMarkdown": "Thanks for promoting pydegensac!",
      "votes": 5
    },
    {
      "id": 1035982,
      "postDate": "2020-10-03T10:04:30.383Z",
      "content": "<p>Thank you for providing the baseline!</p>",
      "rawMarkdown": "Thank you for providing the baseline!"
    },
    {
      "id": 980141,
      "postDate": "2020-08-21T10:56:31.690Z",
      "content": "<p>Thank you for providing the notebook :)</p>",
      "rawMarkdown": "Thank you for providing the notebook :)"
    }
  ],
  "comments": [
    {
      "id": 962022,
      "author_name": "samuel natanael",
      "author_url": "",
      "post_date": "2020-08-07T18:09:55.570000",
      "content": "<p>it threw an exception when i submit, any idea why?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 969364,
          "author_name": "AviralJain",
          "author_url": "",
          "post_date": "2020-08-13T16:26:15.717000",
          "content": "<p>Hey, same happened to me… did you find out the reason?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 976701,
          "author_name": "AmorfEvo",
          "author_url": "",
          "post_date": "2020-08-19T03:38:58.213000",
          "content": "<p>Same - this base kernel just give an error…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 977249,
          "author_name": "omarsamir",
          "author_url": "",
          "post_date": "2020-08-19T11:25:45.523000",
          "content": "<p>I face the same issue here</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 999308,
          "author_name": "omarsamir",
          "author_url": "",
          "post_date": "2020-09-05T14:54:36.807000",
          "content": "<p>I am able to submit successfully several days ago using this kernel (version 9) <a href=\"https://www.kaggle.com/paulorzp/baseline-landmark-recognition-lb-0-48?scriptVersionId=41322107\" target=\"_blank\">https://www.kaggle.com/paulorzp/baseline-landmark-recognition-lb-0-48?scriptVersionId=41322107</a> I think it is only differ from the baseline in just few parameter values. I think also submission failure is due to time exceeding of the kernel.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1011206,
      "author_name": "ayush",
      "author_url": "",
      "post_date": "2020-09-15T10:03:53.733000",
      "content": "<p>Thanks for providing the notebook. It was really very helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 979524,
      "author_name": "omarsamir",
      "author_url": "",
      "post_date": "2020-08-20T23:10:19.573000",
      "content": "<p>I run the base line notebook but I got a submission score error, I see also some comments reporting errors while running the base line notebook.  Did anyone able to run the base line successfully ?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1003358,
      "author_name": "temk",
      "author_url": "",
      "post_date": "2020-09-08T21:46:29.037000",
      "content": "<p>Thank you for the example. <br>\nIn the public train data set, every landmark presented at least twice. Is it correct also for private train set?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 993939,
      "author_name": "Rahul Sattar",
      "author_url": "",
      "post_date": "2020-09-01T08:43:44.357000",
      "content": "<p>Please help me with the following error. </p>\n<p>OSError: SavedModel file does not exist at: ../input/delg-saved-models/local_and_global/{saved_model.pbtxt|saved_model.pb}</p>",
      "votes": 0,
      "replies": [
        {
          "id": 996895,
          "author_name": "Fais K",
          "author_url": "",
          "post_date": "2020-09-03T16:18:25.217000",
          "content": "<p>Try correcting the filename with the file format extension.<br>\nThanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 965316,
      "author_name": "Tolga",
      "author_url": "",
      "post_date": "2020-08-10T14:42:55.593000",
      "content": "<p>I was wondering if the baseline model given in your example an exact copy of the DELG pre-trained on the Google-Landmarks dataset v1 or the RN101-ArcFace pre-trained on Google-Landmarks dataset v2 (clean)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 966568,
          "author_name": "Cam Askew",
          "author_url": "",
          "post_date": "2020-08-11T14:28:48.330000",
          "content": "<p>Yep, the weights are identical to the publicly available pre-trained model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 961054,
      "author_name": "Usmann Khan",
      "author_url": "",
      "post_date": "2020-08-06T21:55:10.480000",
      "content": "<p>Will you be releasing the training code?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 961058,
          "author_name": "Cam Askew",
          "author_url": "",
          "post_date": "2020-08-06T22:04:14.900000",
          "content": "<p>Take a look at <a href=\"https://github.com/tensorflow/models/tree/master/research/delf\">the github repo</a>!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 960866,
      "author_name": "Tolga",
      "author_url": "",
      "post_date": "2020-08-06T18:45:51.793000",
      "content": "<p>Thank you for the baseline example! I would like to ask if the ~10h runtime is achieved on gpu?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 960963,
          "author_name": "Cam Askew",
          "author_url": "",
          "post_date": "2020-08-06T20:24:00.357000",
          "content": "<p>yup using GPU</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 960864,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-08-06T18:42:37.173000",
      "content": "<p>Thanks for providing this notebook! Do you have any resources to get started with image retrieval? 👀 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1009947,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-14T11:48:22.867000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 999185,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-05T12:58:27.160000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 998081,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-04T13:30:48.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 976323,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T19:36:27.083000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 963859,
      "author_name": "old-ufo",
      "author_url": "",
      "post_date": "2020-08-09T11:17:51.240000",
      "content": "<p>Thanks for promoting pydegensac!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1035982,
      "author_name": "Lenscloth",
      "author_url": "",
      "post_date": "2020-10-03T10:04:30.383000",
      "content": "<p>Thank you for providing the baseline!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 980141,
      "author_name": "Ranitra Das",
      "author_url": "",
      "post_date": "2020-08-21T10:56:31.690000",
      "content": "<p>Thank you for providing the notebook :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "959487": "Howdy, folks!\n\nSince this is a code competition, you'll be submitting a Notebook instead of a csv, as in previous years. This means you can use techniques like ensembles and local feature re-ranking and whatever implementation you like, so long as your kernel runs within the hardware constraints of the Notebook worker and finishes in 12h.\n\nThe format *can* make that first submission a little bit more difficult though, so I put together an [example Notebook](https://www.kaggle.com/camaskew/host-baseline-example) and shared it publicly to get you started. There's a more detailed description in the Notebook itself, but in short, it uses TensorFlow to load pre-trained [DELG](https://arxiv.org/abs/2001.05027) models from an attached dataset for feature extraction. It makes predictions using retrieval by global feature similarity and re-ranking by geometric verification.\n\nTo stay below the time limit, I only extracted local features for 3 scales, and only re-ranked the top `NUM_TO_RERANK` similarity-ranked training images, using [pydegensac](https://pypi.org/project/pydegensac/) for speedy geometric-verification. With `NUM_TO_RERANK = 10`, runtime is just under 10h. As a heads-up: subsequent re-runs may produce slightly different results, since there's not an easy way to specify a seed for RANSAC.\nThe code also has a few lines that simplified my submission workflow, and I hope they can help you, too! The dataset for the competition is split into public and private subsets. When your code runs in a Draft Session or a Commit, you'll be given the public images to test and train with. Submissions are run against the private set, though, and since you start a submission by clicking \"Submit\" after a Commit outputs `submission.csv`, the example quickly creates a submission by copying `sample_submission.csv` from the competition dataset. This way, you don't have to run your code on both the private and public datasets to see your actual score.\n\nI hope the example ends up being helpful, and I can't wait to see your improvements on it!\n\nHappy coding,\nCam\n",
    "962022": "it threw an exception when i submit, any idea why?",
    "1011206": "Thanks for providing the notebook. It was really very helpful!",
    "979524": "I run the base line notebook but I got a submission score error, I see also some comments reporting errors while running the base line notebook.  Did anyone able to run the base line successfully ?",
    "1003358": "Thank you for the example. \nIn the public train data set, every landmark presented at least twice. Is it correct also for private train set?",
    "993939": "Please help me with the following error. \n\nOSError: SavedModel file does not exist at: ../input/delg-saved-models/local_and_global/{saved_model.pbtxt|saved_model.pb}",
    "965316": "I was wondering if the baseline model given in your example an exact copy of the DELG pre-trained on the Google-Landmarks dataset v1 or the RN101-ArcFace pre-trained on Google-Landmarks dataset v2 (clean)?",
    "961054": "Will you be releasing the training code?",
    "960866": "Thank you for the baseline example! I would like to ask if the ~10h runtime is achieved on gpu?",
    "960864": "Thanks for providing this notebook! Do you have any resources to get started with image retrieval? 👀 ",
    "1009947": "",
    "999185": "",
    "998081": "",
    "976323": "",
    "963859": "Thanks for promoting pydegensac!",
    "1035982": "Thank you for providing the baseline!",
    "980141": "Thank you for providing the notebook :)"
  }
}