{
  "id": 173091,
  "title": "Some insights",
  "url": "/competitions/landmark-recognition-2020/discussion/173091",
  "author_name": "",
  "post_date": "2020-08-07T19:11:11.091592500Z",
  "votes": 40,
  "comment_count": 27,
  "views": 0,
  "content": "<p>As usual, here are some insights to get started with this competition: </p>\n<ul>\n<li><p>Train images:</p>\n<ul>\n<li>Import the train landmark labels dataset:</li></ul>\n<pre><code>import pandas as pd\ndf = pd.read_csv(\"path/to/train.csv\")\n</code></pre>\n<ul>\n<li>Data comes from this dataset: <a href=\"https://github.com/cvdfoundation/google-landmark\" target=\"_blank\">Google Landmarks Dataset v2</a></li>\n<li>This dataset contains both natural and human-made landmarks.</li>\n<li>You can learn more about the dataset in this <a href=\"https://ai.googleblog.com/2019/05/announcing-google-landmarks-v2-improved.html\" target=\"_blank\">blog post</a>.</li>\n<li>It can be used for both image recognition and image retrieval.</li>\n<li>As you have guessed it, this competition is only focused on image recognition. There is a separate one for <br>\nimage retrieval <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020\" target=\"_blank\">here</a></li>\n<li>There are about <strong>1.5M</strong> images (<code>1_580_470</code> to be precise): <code>len(df)</code></li>\n<li>There are <strong>81313</strong> unique landmark ids: <code>df[\"landmark_id\"].nunique()</code></li>\n<li>Some landmarks have a lot of images (landmard <code>138982</code> has <strong>6272</strong>) where some have very few (landmark <code>197219</code> has <strong>2</strong>). How to deal with these cases?</li>\n<li>Here is the histogram of number of images per landmark where I have dropped the <strong>800</strong> top ones so that the plot isn't too much skewed: </li></ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F57a8ae19798b38eb84fa2336e79f4594%2Flandmark_values_histo.png?generation=1596997870261719&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Here is the code: </li></ul>\n<pre><code>s = df[\"landmark_id\"].value_counts(ascending=True)\nfor _ in range(800):\n    s = s.drop(labels=s.idxmax())\ns.plot(kind=\"hist\", bins=100)\n</code></pre>\n<ul>\n<li><p>As you can see from the histogram, a lot of landmarks have few images (4, 5, 6, and so on) so data augmentation will be useful to help with generalization.</p></li>\n<li><p>Let's explore the images associated with one landmark. For example <code>116375</code>: <code>df.loc[df[\"landmark_id\"] == 116375, \"id\"].tolist()</code></p></li>\n<li><p>You should get the following list: </p></li></ul>\n<pre><code>['36996044c3fdebda',\n'440e872e6216bec5',\n'7334fe9cf5c6f487',\n'9036cc9c516ccea3',\n'a43d62b09cb0a621']\n</code></pre>\n<ul>\n<li>Here is a short code snippet to display these images: </li></ul>\n<pre><code>from pathlib import Path\nlandmark_id = 116375\nimages = df.loc[df[\"landmark_id\"] == landmark_id, \"id\"].tolist()\nbase_folder = \"path/to/train/images\"\nfor image in images:\n    image_folder = \"/\".join(c for c in image[:3])\n    path = Path(base_folder) / image_folder / f\"{image}.jpg\"\n    Image.open(path).show()\n</code></pre>\n<ul>\n<li>Here are the 5 images into a single screenshot: </li></ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F943aefcf556b2ec2c62a19708b2effea%2F116375_images_montage.png?generation=1597000275818884&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Looks like an amusement park. :) </li></ul></li>\n\n\n<li><p>Test images:</p>\n<ul>\n<li>There are <code>10345</code> images in the public test dataset. Before that, I was using the following command: <code>find .//. ! -name . -print | grep -c //</code>. This is almost true but includes the number of folders as well. </li></ul></li>\n<li><p>Evaluation metric</p>\n<ul>\n<li><p>Global Average Precision (GAP) at 1</p></li>\n<li><p>To compute it: </p>\n<ol>\n<li>Predict one <strong>landmark label</strong> for each test image and a <strong>confidence score</strong></li>\n<li>Sort the list of all the test predictions from highest confidence score to lowest.</li>\n<li>Compute the average precision of the above list</li></ol></li>\n<li><p>The formula is: $$GAP = \\frac{1}{M}\\sum_{i=1}^N P(i) rel(i)$$</p></li></ul></li>\n<li><p>Useful concepts:</p>\n<ul>\n<li>Image embeddings</li>\n<li>Distance between two images</li>\n<li>Re-ranking</li></ul></li>\n<li><p>Useful models:</p>\n<ul>\n<li><strong>DELG</strong>: <a href=\"https://arxiv.org/abs/2001.05027\" target=\"_blank\"><strong>paper</strong></a> and <a href=\"https://paperswithcode.com/paper/unifying-deep-local-and-global-features-for\" target=\"_blank\"><strong>implementation</strong></a></li></ul></li>\n<li><p>Some resources:</p>\n<ul>\n<li>The GLDv2 <a href=\"https://github.com/cvdfoundation/google-landmark\" target=\"_blank\">github repo</a></li>\n<li>Landmark workshop <a href=\"https://landmarksworkshop.github.io/CVPRW2019/\" target=\"_blank\">2019 edition</a></li>\n<li>To  learn more about GAP: <a href=\"https://www.researchgate.net/publication/224579197_A_family_of_contextual_measures_of_similarity_between_distributions_with_application_to_image_retrieval\" target=\"_blank\">https://www.researchgate.net/publication/224579197_A_family_of_contextual_measures_of_similarity_between_distributions_with_application_to_image_retrieval</a> </li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "962069",
      "postDate": "08/07/2020 19:11:11",
      "content": "<p>As usual, here are some insights to get started with this competition: </p>\n<ul>\n<li><p>Train images:</p>\n<ul>\n<li>Import the train landmark labels dataset:</li></ul>\n<pre><code>import pandas as pd\ndf = pd.read_csv(\"path/to/train.csv\")\n</code></pre>\n<ul>\n<li>Data comes from this dataset: <a href=\"https://github.com/cvdfoundation/google-landmark\" target=\"_blank\">Google Landmarks Dataset v2</a></li>\n<li>This dataset contains both natural and human-made landmarks.</li>\n<li>You can learn more about the dataset in this <a href=\"https://ai.googleblog.com/2019/05/announcing-google-landmarks-v2-improved.html\" target=\"_blank\">blog post</a>.</li>\n<li>It can be used for both image recognition and image retrieval.</li>\n<li>As you have guessed it, this competition is only focused on image recognition. There is a separate one for <br>\nimage retrieval <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020\" target=\"_blank\">here</a></li>\n<li>There are about <strong>1.5M</strong> images (<code>1_580_470</code> to be precise): <code>len(df)</code></li>\n<li>There are <strong>81313</strong> unique landmark ids: <code>df[\"landmark_id\"].nunique()</code></li>\n<li>Some landmarks have a lot of images (landmard <code>138982</code> has <strong>6272</strong>) where some have very few (landmark <code>197219</code> has <strong>2</strong>). How to deal with these cases?</li>\n<li>Here is the histogram of number of images per landmark where I have dropped the <strong>800</strong> top ones so that the plot isn't too much skewed: </li></ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F57a8ae19798b38eb84fa2336e79f4594%2Flandmark_values_histo.png?generation=1596997870261719&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Here is the code: </li></ul>\n<pre><code>s = df[\"landmark_id\"].value_counts(ascending=True)\nfor _ in range(800):\n    s = s.drop(labels=s.idxmax())\ns.plot(kind=\"hist\", bins=100)\n</code></pre>\n<ul>\n<li><p>As you can see from the histogram, a lot of landmarks have few images (4, 5, 6, and so on) so data augmentation will be useful to help with generalization.</p></li>\n<li><p>Let's explore the images associated with one landmark. For example <code>116375</code>: <code>df.loc[df[\"landmark_id\"] == 116375, \"id\"].tolist()</code></p></li>\n<li><p>You should get the following list: </p></li></ul>\n<pre><code>['36996044c3fdebda',\n'440e872e6216bec5',\n'7334fe9cf5c6f487',\n'9036cc9c516ccea3',\n'a43d62b09cb0a621']\n</code></pre>\n<ul>\n<li>Here is a short code snippet to display these images: </li></ul>\n<pre><code>from pathlib import Path\nlandmark_id = 116375\nimages = df.loc[df[\"landmark_id\"] == landmark_id, \"id\"].tolist()\nbase_folder = \"path/to/train/images\"\nfor image in images:\n    image_folder = \"/\".join(c for c in image[:3])\n    path = Path(base_folder) / image_folder / f\"{image}.jpg\"\n    Image.open(path).show()\n</code></pre>\n<ul>\n<li>Here are the 5 images into a single screenshot: </li></ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F943aefcf556b2ec2c62a19708b2effea%2F116375_images_montage.png?generation=1597000275818884&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Looks like an amusement park. :) </li></ul></li>\n\n\n<li><p>Test images:</p>\n<ul>\n<li>There are <code>10345</code> images in the public test dataset. Before that, I was using the following command: <code>find .//. ! -name . -print | grep -c //</code>. This is almost true but includes the number of folders as well. </li></ul></li>\n<li><p>Evaluation metric</p>\n<ul>\n<li><p>Global Average Precision (GAP) at 1</p></li>\n<li><p>To compute it: </p>\n<ol>\n<li>Predict one <strong>landmark label</strong> for each test image and a <strong>confidence score</strong></li>\n<li>Sort the list of all the test predictions from highest confidence score to lowest.</li>\n<li>Compute the average precision of the above list</li></ol></li>\n<li><p>The formula is: $$GAP = \\frac{1}{M}\\sum_{i=1}^N P(i) rel(i)$$</p></li></ul></li>\n<li><p>Useful concepts:</p>\n<ul>\n<li>Image embeddings</li>\n<li>Distance between two images</li>\n<li>Re-ranking</li></ul></li>\n<li><p>Useful models:</p>\n<ul>\n<li><strong>DELG</strong>: <a href=\"https://arxiv.org/abs/2001.05027\" target=\"_blank\"><strong>paper</strong></a> and <a href=\"https://paperswithcode.com/paper/unifying-deep-local-and-global-features-for\" target=\"_blank\"><strong>implementation</strong></a></li></ul></li>\n<li><p>Some resources:</p>\n<ul>\n<li>The GLDv2 <a href=\"https://github.com/cvdfoundation/google-landmark\" target=\"_blank\">github repo</a></li>\n<li>Landmark workshop <a href=\"https://landmarksworkshop.github.io/CVPRW2019/\" target=\"_blank\">2019 edition</a></li>\n<li>To  learn more about GAP: <a href=\"https://www.researchgate.net/publication/224579197_A_family_of_contextual_measures_of_similarity_between_distributions_with_application_to_image_retrieval\" target=\"_blank\">https://www.researchgate.net/publication/224579197_A_family_of_contextual_measures_of_similarity_between_distributions_with_application_to_image_retrieval</a> </li></ul></li>\n</ul>",
      "rawMarkdown": "As usual, here are some insights to get started with this competition: \n\n\n\n* Train images:\n\n    * Import the train landmark labels dataset:\n\n    ```\n    import pandas as pd\n    df = pd.read_csv(\"path/to/train.csv\")\n    ```  \n\n    * Data comes from this dataset: [Google Landmarks Dataset v2](https://github.com/cvdfoundation/google-landmark)\n    * This dataset contains both natural and human-made landmarks.\n    * You can learn more about the dataset in this [blog post](https://ai.googleblog.com/2019/05/announcing-google-landmarks-v2-improved.html).\n    * It can be used for both image recognition and image retrieval.\n    * As you have guessed it, this competition is only focused on image recognition. There is a separate one for \n    image retrieval [here](https://www.kaggle.com/c/landmark-retrieval-2020)\n    * There are about **1.5M** images (`1_580_470` to be precise): `len(df)`\n    * There are **81313** unique landmark ids: `df[\"landmark_id\"].nunique()`\n    * Some landmarks have a lot of images (landmard `138982` has **6272**) where some have very few (landmark `197219` has **2**). How to deal with these cases?\n    * Here is the histogram of number of images per landmark where I have dropped the **800** top ones so that the plot isn't too much skewed: \n\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F57a8ae19798b38eb84fa2336e79f4594%2Flandmark_values_histo.png?generation=1596997870261719&amp;alt=media)\n\n\n    * Here is the code: \n\n    ```\n    s = df[\"landmark_id\"].value_counts(ascending=True)\n    for _ in range(800):\n        s = s.drop(labels=s.idxmax())\n    s.plot(kind=\"hist\", bins=100)\n    ```\n\n    * As you can see from the histogram, a lot of landmarks have few images (4, 5, 6, and so on) so data augmentation will be useful to help with generalization.\n\n    * Let's explore the images associated with one landmark. For example `116375`: `df.loc[df[\"landmark_id\"] == 116375, \"id\"].tolist()`\n\n\n    * You should get the following list: \n\n    ```\n    ['36996044c3fdebda',\n    '440e872e6216bec5',\n    '7334fe9cf5c6f487',\n    '9036cc9c516ccea3',\n    'a43d62b09cb0a621']\n    ```\n\n    * Here is a short code snippet to display these images: \n\n    ```\n    from pathlib import Path\n    landmark_id = 116375\n    images = df.loc[df[\"landmark_id\"] == landmark_id, \"id\"].tolist()\n    base_folder = \"path/to/train/images\"\n    for image in images:\n        image_folder = \"/\".join(c for c in image[:3])\n        path = Path(base_folder) / image_folder / f\"{image}.jpg\"\n        Image.open(path).show()\n    ```\n\n    * Here are the 5 images into a single screenshot: \n\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F943aefcf556b2ec2c62a19708b2effea%2F116375_images_montage.png?generation=1597000275818884&amp;alt=media)\n\n    * Looks like an amusement park. :) \n\n\n\n\n* Test images:\n\n    * There are `10345 ` images in the public test dataset. Before that, I was using the following command: `find .//. ! -name . -print | grep -c //`. This is almost true but includes the number of folders as well. \n\n\n* Evaluation metric\n\n    * Global Average Precision (GAP) at 1\n    * To compute it: \n        1. Predict one **landmark label** for each test image and a **confidence score**\n        2. Sort the list of all the test predictions from highest confidence score to lowest.\n        3. Compute the average precision of the above list\n\n    * The formula is: $$GAP = \\frac{1}{M}\\sum_{i=1}^N P(i) rel(i)$$\n\n* Useful concepts:\n\n    * Image embeddings\n    * Distance between two images\n    * Re-ranking\n\n\n* Useful models:\n\n    * **DELG**: [**paper**](https://arxiv.org/abs/2001.05027) and [**implementation**](https://paperswithcode.com/paper/unifying-deep-local-and-global-features-for)\n\n\n* Some resources:\n\n    * The GLDv2 [github repo](https://github.com/cvdfoundation/google-landmark)\n    * Landmark workshop [2019 edition](https://landmarksworkshop.github.io/CVPRW2019/)\n    * To  learn more about GAP: https://www.researchgate.net/publication/224579197_A_family_of_contextual_measures_of_similarity_between_distributions_with_application_to_image_retrieval",
      "votes": null
    },
    {
      "id": "969667",
      "postDate": "08/13/2020 20:50:47",
      "content": "<p>Quick question: what is the best way to display a LaTeX formula in a markdown document? <br>\nSo far, here are the two solutions I came across: </p>\n<ul>\n<li>Use an external service that will generate an image through a URL</li>\n<li>Take a screenshot of the rendered formula or use its SVG format</li>\n</ul>\n<p>Other options?</p>",
      "rawMarkdown": "Quick question: what is the best way to display a LaTeX formula in a markdown document? \nSo far, here are the two solutions I came across: \n\n- Use an external service that will generate an image through a URL\n- Take a screenshot of the rendered formula or use its SVG format\n\nOther options?",
      "votes": null
    },
    {
      "id": "969983",
      "postDate": "08/14/2020 05:24:36",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "970006",
      "postDate": "08/14/2020 05:50:39",
      "content": "<p>You are welcome! Let me know if you need more details!</p>",
      "rawMarkdown": "You are welcome! Let me know if you need more details!",
      "votes": null
    },
    {
      "id": "970980",
      "postDate": "08/15/2020 04:16:12",
      "content": "<p>$\\alpha$ is written in Latex.</p>",
      "rawMarkdown": "$\\alpha$ is written in Latex.",
      "votes": null
    },
    {
      "id": "970981",
      "postDate": "08/15/2020 04:17:03",
      "content": "<p>Usually i have tried before that in order to write Latex in HTML, we place the syntax between $$. It doesn't seems to work here on Kaggle. </p>",
      "rawMarkdown": "Usually i have tried before that in order to write Latex in HTML, we place the syntax between $$. It doesn't seems to work here on Kaggle.",
      "votes": null
    },
    {
      "id": "971119",
      "postDate": "08/15/2020 07:37:09",
      "content": "<p>There's a problem displaying Latex formula on Google Chrome. Try using Firefox to see the correct formula displayed</p>",
      "rawMarkdown": "There's a problem displaying Latex formula on Google Chrome. Try using Firefox to see the correct formula displayed",
      "votes": null
    },
    {
      "id": "971643",
      "postDate": "08/15/2020 18:51:58",
      "content": "<p><a href=\"https://www.kaggle.com/jamshaidsohail5\" target=\"_blank\">@jamshaidsohail5</a> I guess you are referring to Kaggle notebooks. If that's the case, we can indeed place Latex commands there using what you have mentioned. As far as I know, this isn't possible natively in Markdown. I will keep investigating and update with the best solution. </p>",
      "rawMarkdown": "jamshaidsohail5 I guess you are referring to Kaggle notebooks. If that's the case, we can indeed place Latex commands there using what you have mentioned. As far as I know, this isn't possible natively in Markdown. I will keep investigating and update with the best solution.",
      "votes": null
    },
    {
      "id": "971644",
      "postDate": "08/15/2020 18:53:40",
      "content": "<p>Latest test: $$\\frac{1}{2}$$ does work (i.e <code>$$</code> before and after the equation). Thanks for the pointer <a href=\"https://www.kaggle.com/jamshaidsohail5\" target=\"_blank\">@jamshaidsohail5</a>. 👍👍👍</p>",
      "rawMarkdown": "Latest test: $$\\frac{1}{2}$$ does work (i.e `$$` before and after the equation). Thanks for the pointer @jamshaidsohail5. 👍👍👍",
      "votes": null
    },
    {
      "id": "971645",
      "postDate": "08/15/2020 18:55:46",
      "content": "<p>This means that Kaggle  commenting form does some <a href=\"https://www.mathjax.org/\" target=\"_blank\"><strong>MathJax</strong></a> (or other flavors) processing. 👌</p>\n<p>Update: it is indeed MathJax as displayed in this screenshot</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F8401fe0e4abe0dc0ee7a42ad15091c07%2Fmathjax.png?generation=1597518082360358&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This means that Kaggle  commenting form does some [**MathJax**](https://www.mathjax.org/) (or other flavors) processing. 👌\n\nUpdate: it is indeed MathJax as displayed in this screenshot\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F8401fe0e4abe0dc0ee7a42ad15091c07%2Fmathjax.png?generation=1597518082360358&alt=media)",
      "votes": null
    },
    {
      "id": "973799",
      "postDate": "08/17/2020 14:20:49",
      "content": "<p>Hi, thanks for sharing this, really helpful. Two things:<br>\n1) Number of test images I got was 10345 (len([x for x in pathlib.Path(image_root_dir).rglob('*.jpg')])). Are you sure you're not counting the folders also in your command?<br>\n2) Do we have any idea what is the train-test split ratio in the private data set. Is the private test set same as the public test set?</p>",
      "rawMarkdown": "Hi, thanks for sharing this, really helpful. Two things:\n1) Number of test images I got was 10345 (len([x for x in pathlib.Path(image_root_dir).rglob('*.jpg')])). Are you sure you're not counting the folders also in your command?\n2) Do we have any idea what is the train-test split ratio in the private data set. Is the private test set same as the public test set?",
      "votes": null
    },
    {
      "id": "974188",
      "postDate": "08/17/2020 19:28:28",
      "content": "<p>Thanks for these remarks:</p>\n<ol>\n<li>I will double check and let you know.</li>\n<li>If you check the leaderboard tab, you will get the split:</li>\n</ol>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 34% of the test data.</p>\n  <p>The final results will be based on the other 66%, so the final standings may be different. </p>\n</blockquote>\n<p>As for the second part of the question, it won't be the same given the splits. Plus, I don't think there will be any overlap but I don't have any data to back this up for now.</p>\n<p>Best of luck!</p>",
      "rawMarkdown": "Thanks for these remarks:\n\n1. I will double check and let you know.\n2. If you check the leaderboard tab, you will get the split:\n\n> This leaderboard is calculated with approximately 34% of the test data.\n\n> The final results will be based on the other 66%, so the final standings may be different. \n\nAs for the second part of the question, it won't be the same given the splits. Plus, I don't think there will be any overlap but I don't have any data to back this up for now.\n\nBest of luck!",
      "votes": null
    },
    {
      "id": "974743",
      "postDate": "08/18/2020 03:21:34",
      "content": "<p>In the second question, i don't mean the private and public leaderboards. The format of this competition is a bit different, in that the whole test and training sets provided in the synchronous run are private and different from what is available publicly. </p>",
      "rawMarkdown": "In the second question, i don't mean the private and public leaderboards. The format of this competition is a bit different, in that the whole test and training sets provided in the synchronous run are private and different from what is available publicly.",
      "votes": null
    },
    {
      "id": "977880",
      "postDate": "08/19/2020 19:12:09",
      "content": "<p>Alright, I see. Since this is a landmark recognition task, all test images are included in the train dateset as mentioned here:</p>\n<blockquote>\n  <p>To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set. You may still attach the full training set as an external data set if you wish.</p>\n</blockquote>\n<p>I am not sure I understand this 100%. I will update the thread if I have more details.</p>",
      "rawMarkdown": "Alright, I see. Since this is a landmark recognition task, all test images are included in the train dateset as mentioned here:\n\n> To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set. You may still attach the full training set as an external data set if you wish.\n\nI am not sure I understand this 100%. I will update the thread if I have more details.",
      "votes": null
    },
    {
      "id": "978620",
      "postDate": "08/20/2020 09:39:12",
      "content": "<p>Great! Thanks for sharing!</p>",
      "rawMarkdown": "Great! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "978626",
      "postDate": "08/20/2020 09:43:01",
      "content": "<p>Glad it helps. Stay tuned for more content, I tend to update this thread over time. ;)</p>",
      "rawMarkdown": "Glad it helps. Stay tuned for more content, I tend to update this thread over time. ;)",
      "votes": null
    },
    {
      "id": "978665",
      "postDate": "08/20/2020 10:03:13",
      "content": "<p>Looking forward to your next content!</p>",
      "rawMarkdown": "Looking forward to your next content!",
      "votes": null
    },
    {
      "id": "978765",
      "postDate": "08/20/2020 11:20:28",
      "content": "<p>A notebook is coming at some point in the next week or so. Stay tuned. ;)</p>",
      "rawMarkdown": "A notebook is coming at some point in the next week or so. Stay tuned. ;)",
      "votes": null
    },
    {
      "id": "979711",
      "postDate": "08/21/2020 04:34:20",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "980359",
      "postDate": "08/21/2020 14:35:19",
      "content": "<p>My pleasure. Let me know if you need more details for some aspects. </p>",
      "rawMarkdown": "My pleasure. Let me know if you need more details for some aspects.",
      "votes": null
    },
    {
      "id": "980959",
      "postDate": "08/22/2020 03:51:03",
      "content": "<p>Public train - 1.5M<br>\nPrivate train - selected 100k, just to reduce repetitive work as you cannot compare with 1.5M within 12hrs.</p>\n<p>If they did not introduce the idea of private train many people would have written code to filter out samples from 1.5M train.csv to retrieve neighbours from.</p>\n<p>Also, it is 10345 public test images, the 14k number came because the command outputs folder names as well</p>",
      "rawMarkdown": "Public train - 1.5M\nPrivate train - selected 100k, just to reduce repetitive work as you cannot compare with 1.5M within 12hrs.\n\nIf they did not introduce the idea of private train many people would have written code to filter out samples from 1.5M train.csv to retrieve neighbours from.\n\nAlso, it is 10345 public test images, the 14k number came because the command outputs folder names as well",
      "votes": null
    },
    {
      "id": "982187",
      "postDate": "08/23/2020 06:07:57",
      "content": "<p>Quick question please regaridng private 100k data - this is to extract embeddings from and then compare using test set? So, that we don't need to compare with the 1.5M images?</p>",
      "rawMarkdown": "Quick question please regaridng private 100k data - this is to extract embeddings from and then compare using test set? So, that we don't need to compare with the 1.5M images?",
      "votes": null
    },
    {
      "id": "982259",
      "postDate": "08/23/2020 07:46:42",
      "content": "<p>Yes. If you want to compare with 1.5M you can add it as a separate input data.</p>",
      "rawMarkdown": "Yes. If you want to compare with 1.5M you can add it as a separate input data.",
      "votes": null
    },
    {
      "id": "982308",
      "postDate": "08/23/2020 08:48:13",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/skrish13\" target=\"_blank\">@skrish13</a> for all the details. I will update the post accordingly. </p>",
      "rawMarkdown": "Thanks @skrish13 for all the details. I will update the post accordingly.",
      "votes": null
    },
    {
      "id": "982573",
      "postDate": "08/23/2020 13:47:08",
      "content": "<p>Thanks, it was really useful.</p>",
      "rawMarkdown": "Thanks, it was really useful.",
      "votes": null
    },
    {
      "id": "983981",
      "postDate": "08/24/2020 18:30:36",
      "content": "<p>Glad it helps!</p>",
      "rawMarkdown": "Glad it helps!",
      "votes": null
    },
    {
      "id": "993620",
      "postDate": "09/01/2020 04:11:04",
      "content": "<p>Thanks for sharing these insights and links to relevant resources. Very helpful!</p>",
      "rawMarkdown": "Thanks for sharing these insights and links to relevant resources. Very helpful!",
      "votes": null
    },
    {
      "id": "1298212",
      "postDate": "05/08/2021 16:48:44",
      "content": "<p>I am glad it helps. I was planning to add more things but due to lack of time, I didn't. Maybe at some point… :D</p>",
      "rawMarkdown": "I am glad it helps. I was planning to add more things but due to lack of time, I didn't. Maybe at some point... :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969667,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "08/13/2020 20:50:47",
      "content": "<p>Quick question: what is the best way to display a LaTeX formula in a markdown document? <br>\nSo far, here are the two solutions I came across: </p>\n<ul>\n<li>Use an external service that will generate an image through a URL</li>\n<li>Take a screenshot of the rendered formula or use its SVG format</li>\n</ul>\n<p>Other options?</p>",
      "votes": null,
      "replies": [
        {
          "id": 970980,
          "author_name": "jamshaidsohail5",
          "author_url": "",
          "post_date": "08/15/2020 04:16:12",
          "content": "<p>$\\alpha$ is written in Latex.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970981,
          "author_name": "jamshaidsohail5",
          "author_url": "",
          "post_date": "08/15/2020 04:17:03",
          "content": "<p>Usually i have tried before that in order to write Latex in HTML, we place the syntax between $$. It doesn't seems to work here on Kaggle. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971119,
          "author_name": "sidneyng",
          "author_url": "",
          "post_date": "08/15/2020 07:37:09",
          "content": "<p>There's a problem displaying Latex formula on Google Chrome. Try using Firefox to see the correct formula displayed</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971643,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/15/2020 18:51:58",
          "content": "<p><a href=\"https://www.kaggle.com/jamshaidsohail5\" target=\"_blank\">@jamshaidsohail5</a> I guess you are referring to Kaggle notebooks. If that's the case, we can indeed place Latex commands there using what you have mentioned. As far as I know, this isn't possible natively in Markdown. I will keep investigating and update with the best solution. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971644,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/15/2020 18:53:40",
          "content": "<p>Latest test: $$\\frac{1}{2}$$ does work (i.e <code>$$</code> before and after the equation). Thanks for the pointer <a href=\"https://www.kaggle.com/jamshaidsohail5\" target=\"_blank\">@jamshaidsohail5</a>. 👍👍👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971645,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/15/2020 18:55:46",
          "content": "<p>This means that Kaggle  commenting form does some <a href=\"https://www.mathjax.org/\" target=\"_blank\"><strong>MathJax</strong></a> (or other flavors) processing. 👌</p>\n<p>Update: it is indeed MathJax as displayed in this screenshot</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F8401fe0e4abe0dc0ee7a42ad15091c07%2Fmathjax.png?generation=1597518082360358&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 969983,
      "author_name": "amanrajput27",
      "author_url": "",
      "post_date": "08/14/2020 05:24:36",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 970006,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/14/2020 05:50:39",
          "content": "<p>You are welcome! Let me know if you need more details!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 973799,
      "author_name": "aviraljn9",
      "author_url": "",
      "post_date": "08/17/2020 14:20:49",
      "content": "<p>Hi, thanks for sharing this, really helpful. Two things:<br>\n1) Number of test images I got was 10345 (len([x for x in pathlib.Path(image_root_dir).rglob('*.jpg')])). Are you sure you're not counting the folders also in your command?<br>\n2) Do we have any idea what is the train-test split ratio in the private data set. Is the private test set same as the public test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 974188,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/17/2020 19:28:28",
          "content": "<p>Thanks for these remarks:</p>\n<ol>\n<li>I will double check and let you know.</li>\n<li>If you check the leaderboard tab, you will get the split:</li>\n</ol>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 34% of the test data.</p>\n  <p>The final results will be based on the other 66%, so the final standings may be different. </p>\n</blockquote>\n<p>As for the second part of the question, it won't be the same given the splits. Plus, I don't think there will be any overlap but I don't have any data to back this up for now.</p>\n<p>Best of luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 974743,
          "author_name": "aviraljn9",
          "author_url": "",
          "post_date": "08/18/2020 03:21:34",
          "content": "<p>In the second question, i don't mean the private and public leaderboards. The format of this competition is a bit different, in that the whole test and training sets provided in the synchronous run are private and different from what is available publicly. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 977880,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/19/2020 19:12:09",
          "content": "<p>Alright, I see. Since this is a landmark recognition task, all test images are included in the train dateset as mentioned here:</p>\n<blockquote>\n  <p>To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set. You may still attach the full training set as an external data set if you wish.</p>\n</blockquote>\n<p>I am not sure I understand this 100%. I will update the thread if I have more details.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980959,
          "author_name": "skrish13",
          "author_url": "",
          "post_date": "08/22/2020 03:51:03",
          "content": "<p>Public train - 1.5M<br>\nPrivate train - selected 100k, just to reduce repetitive work as you cannot compare with 1.5M within 12hrs.</p>\n<p>If they did not introduce the idea of private train many people would have written code to filter out samples from 1.5M train.csv to retrieve neighbours from.</p>\n<p>Also, it is 10345 public test images, the 14k number came because the command outputs folder names as well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 982187,
          "author_name": "aroraaman",
          "author_url": "",
          "post_date": "08/23/2020 06:07:57",
          "content": "<p>Quick question please regaridng private 100k data - this is to extract embeddings from and then compare using test set? So, that we don't need to compare with the 1.5M images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 982259,
          "author_name": "skrish13",
          "author_url": "",
          "post_date": "08/23/2020 07:46:42",
          "content": "<p>Yes. If you want to compare with 1.5M you can add it as a separate input data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 982308,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/23/2020 08:48:13",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/skrish13\" target=\"_blank\">@skrish13</a> for all the details. I will update the post accordingly. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 978620,
      "author_name": "salmaneunus",
      "author_url": "",
      "post_date": "08/20/2020 09:39:12",
      "content": "<p>Great! Thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 978626,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/20/2020 09:43:01",
          "content": "<p>Glad it helps. Stay tuned for more content, I tend to update this thread over time. ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 978665,
          "author_name": "salmaneunus",
          "author_url": "",
          "post_date": "08/20/2020 10:03:13",
          "content": "<p>Looking forward to your next content!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 978765,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/20/2020 11:20:28",
          "content": "<p>A notebook is coming at some point in the next week or so. Stay tuned. ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 979711,
      "author_name": "jumpingdino",
      "author_url": "",
      "post_date": "08/21/2020 04:34:20",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 980359,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/21/2020 14:35:19",
          "content": "<p>My pleasure. Let me know if you need more details for some aspects. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 982573,
      "author_name": "vishnus",
      "author_url": "",
      "post_date": "08/23/2020 13:47:08",
      "content": "<p>Thanks, it was really useful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 983981,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "08/24/2020 18:30:36",
          "content": "<p>Glad it helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 993620,
      "author_name": "balraj98",
      "author_url": "",
      "post_date": "09/01/2020 04:11:04",
      "content": "<p>Thanks for sharing these insights and links to relevant resources. Very helpful!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1298212,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "05/08/2021 16:48:44",
          "content": "<p>I am glad it helps. I was planning to add more things but due to lack of time, I didn't. Maybe at some point… :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "962069": "As usual, here are some insights to get started with this competition: \n\n\n\n* Train images:\n\n    * Import the train landmark labels dataset:\n\n    ```\n    import pandas as pd\n    df = pd.read_csv(\"path/to/train.csv\")\n    ```  \n\n    * Data comes from this dataset: [Google Landmarks Dataset v2](https://github.com/cvdfoundation/google-landmark)\n    * This dataset contains both natural and human-made landmarks.\n    * You can learn more about the dataset in this [blog post](https://ai.googleblog.com/2019/05/announcing-google-landmarks-v2-improved.html).\n    * It can be used for both image recognition and image retrieval.\n    * As you have guessed it, this competition is only focused on image recognition. There is a separate one for \n    image retrieval [here](https://www.kaggle.com/c/landmark-retrieval-2020)\n    * There are about **1.5M** images (`1_580_470` to be precise): `len(df)`\n    * There are **81313** unique landmark ids: `df[\"landmark_id\"].nunique()`\n    * Some landmarks have a lot of images (landmard `138982` has **6272**) where some have very few (landmark `197219` has **2**). How to deal with these cases?\n    * Here is the histogram of number of images per landmark where I have dropped the **800** top ones so that the plot isn't too much skewed: \n\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F57a8ae19798b38eb84fa2336e79f4594%2Flandmark_values_histo.png?generation=1596997870261719&amp;alt=media)\n\n\n    * Here is the code: \n\n    ```\n    s = df[\"landmark_id\"].value_counts(ascending=True)\n    for _ in range(800):\n        s = s.drop(labels=s.idxmax())\n    s.plot(kind=\"hist\", bins=100)\n    ```\n\n    * As you can see from the histogram, a lot of landmarks have few images (4, 5, 6, and so on) so data augmentation will be useful to help with generalization.\n\n    * Let's explore the images associated with one landmark. For example `116375`: `df.loc[df[\"landmark_id\"] == 116375, \"id\"].tolist()`\n\n\n    * You should get the following list: \n\n    ```\n    ['36996044c3fdebda',\n    '440e872e6216bec5',\n    '7334fe9cf5c6f487',\n    '9036cc9c516ccea3',\n    'a43d62b09cb0a621']\n    ```\n\n    * Here is a short code snippet to display these images: \n\n    ```\n    from pathlib import Path\n    landmark_id = 116375\n    images = df.loc[df[\"landmark_id\"] == landmark_id, \"id\"].tolist()\n    base_folder = \"path/to/train/images\"\n    for image in images:\n        image_folder = \"/\".join(c for c in image[:3])\n        path = Path(base_folder) / image_folder / f\"{image}.jpg\"\n        Image.open(path).show()\n    ```\n\n    * Here are the 5 images into a single screenshot: \n\n    ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F172860%2F943aefcf556b2ec2c62a19708b2effea%2F116375_images_montage.png?generation=1597000275818884&amp;alt=media)\n\n    * Looks like an amusement park. :) \n\n\n\n\n* Test images:\n\n    * There are `10345 ` images in the public test dataset. Before that, I was using the following command: `find .//. ! -name . -print | grep -c //`. This is almost true but includes the number of folders as well. \n\n\n* Evaluation metric\n\n    * Global Average Precision (GAP) at 1\n    * To compute it: \n        1. Predict one **landmark label** for each test image and a **confidence score**\n        2. Sort the list of all the test predictions from highest confidence score to lowest.\n        3. Compute the average precision of the above list\n\n    * The formula is: $$GAP = \\frac{1}{M}\\sum_{i=1}^N P(i) rel(i)$$\n\n* Useful concepts:\n\n    * Image embeddings\n    * Distance between two images\n    * Re-ranking\n\n\n* Useful models:\n\n    * **DELG**: [**paper**](https://arxiv.org/abs/2001.05027) and [**implementation**](https://paperswithcode.com/paper/unifying-deep-local-and-global-features-for)\n\n\n* Some resources:\n\n    * The GLDv2 [github repo](https://github.com/cvdfoundation/google-landmark)\n    * Landmark workshop [2019 edition](https://landmarksworkshop.github.io/CVPRW2019/)\n    * To  learn more about GAP: https://www.researchgate.net/publication/224579197_A_family_of_contextual_measures_of_similarity_between_distributions_with_application_to_image_retrieval",
    "969667": "Quick question: what is the best way to display a LaTeX formula in a markdown document? \nSo far, here are the two solutions I came across: \n\n- Use an external service that will generate an image through a URL\n- Take a screenshot of the rendered formula or use its SVG format\n\nOther options?",
    "969983": "Thanks for sharing",
    "970006": "You are welcome! Let me know if you need more details!",
    "970980": "$\\alpha$ is written in Latex.",
    "970981": "Usually i have tried before that in order to write Latex in HTML, we place the syntax between $$. It doesn't seems to work here on Kaggle.",
    "971119": "There's a problem displaying Latex formula on Google Chrome. Try using Firefox to see the correct formula displayed",
    "971643": "jamshaidsohail5 I guess you are referring to Kaggle notebooks. If that's the case, we can indeed place Latex commands there using what you have mentioned. As far as I know, this isn't possible natively in Markdown. I will keep investigating and update with the best solution.",
    "971644": "Latest test: $$\\frac{1}{2}$$ does work (i.e `$$` before and after the equation). Thanks for the pointer @jamshaidsohail5. 👍👍👍",
    "971645": "This means that Kaggle  commenting form does some [**MathJax**](https://www.mathjax.org/) (or other flavors) processing. 👌\n\nUpdate: it is indeed MathJax as displayed in this screenshot\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F172860%2F8401fe0e4abe0dc0ee7a42ad15091c07%2Fmathjax.png?generation=1597518082360358&alt=media)",
    "973799": "Hi, thanks for sharing this, really helpful. Two things:\n1) Number of test images I got was 10345 (len([x for x in pathlib.Path(image_root_dir).rglob('*.jpg')])). Are you sure you're not counting the folders also in your command?\n2) Do we have any idea what is the train-test split ratio in the private data set. Is the private test set same as the public test set?",
    "974188": "Thanks for these remarks:\n\n1. I will double check and let you know.\n2. If you check the leaderboard tab, you will get the split:\n\n> This leaderboard is calculated with approximately 34% of the test data.\n\n> The final results will be based on the other 66%, so the final standings may be different. \n\nAs for the second part of the question, it won't be the same given the splits. Plus, I don't think there will be any overlap but I don't have any data to back this up for now.\n\nBest of luck!",
    "974743": "In the second question, i don't mean the private and public leaderboards. The format of this competition is a bit different, in that the whole test and training sets provided in the synchronous run are private and different from what is available publicly.",
    "977880": "Alright, I see. Since this is a landmark recognition task, all test images are included in the train dateset as mentioned here:\n\n> To facilitate recognition-by-retrieval approaches, the private training set contains only a 100k subset of the total public training set. This 100k subset contains all of the training set images associated with the landmarks in the private test set. You may still attach the full training set as an external data set if you wish.\n\nI am not sure I understand this 100%. I will update the thread if I have more details.",
    "978620": "Great! Thanks for sharing!",
    "978626": "Glad it helps. Stay tuned for more content, I tend to update this thread over time. ;)",
    "978665": "Looking forward to your next content!",
    "978765": "A notebook is coming at some point in the next week or so. Stay tuned. ;)",
    "979711": "Thanks for sharing!",
    "980359": "My pleasure. Let me know if you need more details for some aspects.",
    "980959": "Public train - 1.5M\nPrivate train - selected 100k, just to reduce repetitive work as you cannot compare with 1.5M within 12hrs.\n\nIf they did not introduce the idea of private train many people would have written code to filter out samples from 1.5M train.csv to retrieve neighbours from.\n\nAlso, it is 10345 public test images, the 14k number came because the command outputs folder names as well",
    "982187": "Quick question please regaridng private 100k data - this is to extract embeddings from and then compare using test set? So, that we don't need to compare with the 1.5M images?",
    "982259": "Yes. If you want to compare with 1.5M you can add it as a separate input data.",
    "982308": "Thanks @skrish13 for all the details. I will update the post accordingly.",
    "982573": "Thanks, it was really useful.",
    "983981": "Glad it helps!",
    "993620": "Thanks for sharing these insights and links to relevant resources. Very helpful!",
    "1298212": "I am glad it helps. I was planning to add more things but due to lack of time, I didn't. Maybe at some point... :D"
  },
  "source": "meta"
}