{
  "id": 402619,
  "title": "Visually explore the dataset!",
  "url": "/competitions/image-matching-challenge-2023/discussion/402619",
  "author_name": "",
  "post_date": "2023-04-19T03:22:13.719451300Z",
  "votes": 27,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi everyone, I put together a Python script to visualize and explore the dataset for this challenge. Check it out! 😄</p>\n<p>If you find this helpful, consider giving FiftyOne a star on GitHub. It's free and open source!<br>\n<a href=\"https://github.com/voxel51/fiftyone\" target=\"_blank\">https://github.com/voxel51/fiftyone</a></p>\n<h2>Setup</h2>\n<pre><code>pip install fiftyone\n</code></pre>\n<h2>Load dataset</h2>\n<pre><code> glob\n fiftyone  fo\n\n\ndataset_dir = \n\nsamples = []\n filepath  glob.glob(os.path.join(dataset_dir, ), recursive=):\n     filepath.endswith((, , , )):\n        folders = filepath[(dataset_dir) + :].split()[:-]\n        sample = fo.Sample(\n            filepath=filepath,\n            tags=[folders[]],\n            =folders[],\n            location=folders[],\n        )\n        samples.append(sample)\n\ndataset = fo.Dataset(, persistent=)\ndataset.add_samples(samples)\ndataset.compute_metadata()\n</code></pre>\n<h2>Optional: add embeddings</h2>\n<pre><code> fiftyone.brain  fob\n\nfob.compute_visualization(\n    dataset,\n    model=,\n    brain_key=,\n)\n</code></pre>\n<h2>Visualize in the App</h2>\n<pre><code>session = fo.launch_app(dataset)\n</code></pre>\n<h2>Screenshots</h2>\n<p>I could only upload images here, but I uploaded some screen recordings of me exploring the dataset on GitHub here: <a href=\"https://github.com/voxel51/fiftyone/issues/2906\" target=\"_blank\">https://github.com/voxel51/fiftyone/issues/2906</a></p>\n<p><img alt=\"Screen Shot 2023-04-18 at 11 12 30 PM\" src=\"https://user-images.githubusercontent.com/25985824/232957628-091a8e7a-3d45-4a55-8fee-30c366825cac.png\"></p>\n<p><img alt=\"Screen Shot 2023-04-18 at 11 13 30 PM\" src=\"https://user-images.githubusercontent.com/25985824/232957624-f43b7c26-6f65-4a63-b291-d47a7aa07d1f.png\"></p>",
  "messages": [
    {
      "id": "2226556",
      "postDate": "04/19/2023 03:22:13",
      "content": "<p>Hi everyone, I put together a Python script to visualize and explore the dataset for this challenge. Check it out! 😄</p>\n<p>If you find this helpful, consider giving FiftyOne a star on GitHub. It's free and open source!<br>\n<a href=\"https://github.com/voxel51/fiftyone\" target=\"_blank\">https://github.com/voxel51/fiftyone</a></p>\n<h2>Setup</h2>\n<pre><code>pip install fiftyone\n</code></pre>\n<h2>Load dataset</h2>\n<pre><code> glob\n fiftyone  fo\n\n\ndataset_dir = \n\nsamples = []\n filepath  glob.glob(os.path.join(dataset_dir, ), recursive=):\n     filepath.endswith((, , , )):\n        folders = filepath[(dataset_dir) + :].split()[:-]\n        sample = fo.Sample(\n            filepath=filepath,\n            tags=[folders[]],\n            =folders[],\n            location=folders[],\n        )\n        samples.append(sample)\n\ndataset = fo.Dataset(, persistent=)\ndataset.add_samples(samples)\ndataset.compute_metadata()\n</code></pre>\n<h2>Optional: add embeddings</h2>\n<pre><code> fiftyone.brain  fob\n\nfob.compute_visualization(\n    dataset,\n    model=,\n    brain_key=,\n)\n</code></pre>\n<h2>Visualize in the App</h2>\n<pre><code>session = fo.launch_app(dataset)\n</code></pre>\n<h2>Screenshots</h2>\n<p>I could only upload images here, but I uploaded some screen recordings of me exploring the dataset on GitHub here: <a href=\"https://github.com/voxel51/fiftyone/issues/2906\" target=\"_blank\">https://github.com/voxel51/fiftyone/issues/2906</a></p>\n<p><img alt=\"Screen Shot 2023-04-18 at 11 12 30 PM\" src=\"https://user-images.githubusercontent.com/25985824/232957628-091a8e7a-3d45-4a55-8fee-30c366825cac.png\"></p>\n<p><img alt=\"Screen Shot 2023-04-18 at 11 13 30 PM\" src=\"https://user-images.githubusercontent.com/25985824/232957624-f43b7c26-6f65-4a63-b291-d47a7aa07d1f.png\"></p>",
      "rawMarkdown": "Hi everyone, I put together a Python script to visualize and explore the dataset for this challenge. Check it out! 😄\n\nIf you find this helpful, consider giving FiftyOne a star on GitHub. It's free and open source!\nhttps://github.com/voxel51/fiftyone\n\n## Setup\n\n```\npip install fiftyone\n```\n\n## Load dataset\n\n```py\nimport glob\nimport fiftyone as fo\n\n# Download and unzip `image-matching-challenge-2023.zip` and put path here\ndataset_dir = \"/path/to/image-matching-challenge-2023\"\n\nsamples = []\nfor filepath in glob.glob(os.path.join(dataset_dir, \"**\"), recursive=True):\n    if filepath.endswith((\".jpg\", \".jpeg\", \".png\", \".JPG\")):\n        folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]\n        sample = fo.Sample(\n            filepath=filepath,\n            tags=[folders[0]],\n            type=folders[1],\n            location=folders[2],\n        )\n        samples.append(sample)\n\ndataset = fo.Dataset(\"image-matching-challenge-2023\", persistent=True)\ndataset.add_samples(samples)\ndataset.compute_metadata()\n```\n\n## Optional: add embeddings\n\n```py\nimport fiftyone.brain as fob\n\nfob.compute_visualization(\n    dataset,\n    model=\"clip-vit-base32-torch\",\n    brain_key=\"img_viz\",\n)\n```\n\n## Visualize in the App\n\n```py\nsession = fo.launch_app(dataset)\n```\n\n## Screenshots\n\nI could only upload images here, but I uploaded some screen recordings of me exploring the dataset on GitHub here: https://github.com/voxel51/fiftyone/issues/2906\n\n<img width=\"1132\" alt=\"Screen Shot 2023-04-18 at 11 12 30 PM\" src=\"https://user-images.githubusercontent.com/25985824/232957628-091a8e7a-3d45-4a55-8fee-30c366825cac.png\">\n\n<img width=\"2041\" alt=\"Screen Shot 2023-04-18 at 11 13 30 PM\" src=\"https://user-images.githubusercontent.com/25985824/232957624-f43b7c26-6f65-4a63-b291-d47a7aa07d1f.png\">",
      "votes": null
    },
    {
      "id": "2228134",
      "postDate": "04/20/2023 09:41:35",
      "content": "<p>Great work! I put the star <a href=\"https://www.kaggle.com/brimoor\" target=\"_blank\">@brimoor</a></p>",
      "rawMarkdown": "Great work! I put the star @brimoor",
      "votes": null
    },
    {
      "id": "2228691",
      "postDate": "04/20/2023 18:10:40",
      "content": "<p>For detailed instructions on how to get FiftyOne installed, computing the embeddings and further data explorations, check out the full write up here: <a href=\"https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset/\" target=\"_blank\">https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset/</a></p>",
      "rawMarkdown": "For detailed instructions on how to get FiftyOne installed, computing the embeddings and further data explorations, check out the full write up here: https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset/",
      "votes": null
    },
    {
      "id": "2228697",
      "postDate": "04/20/2023 18:12:44",
      "content": "<p>Thanks for the support! 🙌</p>",
      "rawMarkdown": "Thanks for the support! 🙌",
      "votes": null
    },
    {
      "id": "2262572",
      "postDate": "05/17/2023 02:49:24",
      "content": "<p>folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]<br>\nif you find something wrong on you PC,you may need folders = filepath[len(dataset_dir) + 1:].split(\"\\\")[:-2]instead</p>",
      "rawMarkdown": "folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]\nif you find something wrong on you PC,you may need folders = filepath[len(dataset_dir) + 1:].split(\"\\\\\")[:-2]instead",
      "votes": null
    },
    {
      "id": "2262576",
      "postDate": "05/17/2023 02:50:03",
      "content": "<p>folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]<br>\nif you find something wrong on you PC <br>\nyou may need<br>\n split(\"\\\")[:-2]<br>\ninstead</p>",
      "rawMarkdown": "folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]\nif you find something wrong on you PC \nyou may need\n split(\"\\\\\")[:-2]\ninstead",
      "votes": null
    },
    {
      "id": "2262583",
      "postDate": "05/17/2023 02:59:45",
      "content": "<p>how can i open it on an app instead of a jupyter kernel</p>",
      "rawMarkdown": "how can i open it on an app instead of a jupyter kernel",
      "votes": null
    },
    {
      "id": "2262594",
      "postDate": "05/17/2023 03:32:08",
      "content": "<p>When you're working in Jupyter, you can launch the App in a separate tab rather than in the output cell like this:</p>\n<pre><code>\nsession = fo.launch_app(dataset, auto=)\nsession.open_tab()\n</code></pre>\n<p>Documentation: <a href=\"https://docs.voxel51.com/environments/index.html#opening-the-app-in-a-dedicated-tab\" target=\"_blank\">https://docs.voxel51.com/environments/index.html#opening-the-app-in-a-dedicated-tab</a></p>",
      "rawMarkdown": "When you're working in Jupyter, you can launch the App in a separate tab rather than in the output cell like this:\n\n```py\n# Launch the App in a dedicated browser tab\nsession = fo.launch_app(dataset, auto=False)\nsession.open_tab()\n```\n\nDocumentation: https://docs.voxel51.com/environments/index.html#opening-the-app-in-a-dedicated-tab",
      "votes": null
    },
    {
      "id": "2262643",
      "postDate": "05/17/2023 04:26:56",
      "content": "<p>really helpful,thanks!</p>",
      "rawMarkdown": "really helpful,thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2228134,
      "author_name": "yutodennou",
      "author_url": "",
      "post_date": "04/20/2023 09:41:35",
      "content": "<p>Great work! I put the star <a href=\"https://www.kaggle.com/brimoor\" target=\"_blank\">@brimoor</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2228697,
          "author_name": "brimoor",
          "author_url": "",
          "post_date": "04/20/2023 18:12:44",
          "content": "<p>Thanks for the support! 🙌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228691,
      "author_name": "jimmyguerrero999",
      "author_url": "",
      "post_date": "04/20/2023 18:10:40",
      "content": "<p>For detailed instructions on how to get FiftyOne installed, computing the embeddings and further data explorations, check out the full write up here: <a href=\"https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset/\" target=\"_blank\">https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2262572,
      "author_name": "distiller",
      "author_url": "",
      "post_date": "05/17/2023 02:49:24",
      "content": "<p>folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]<br>\nif you find something wrong on you PC,you may need folders = filepath[len(dataset_dir) + 1:].split(\"\\\")[:-2]instead</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2262576,
      "author_name": "distiller",
      "author_url": "",
      "post_date": "05/17/2023 02:50:03",
      "content": "<p>folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]<br>\nif you find something wrong on you PC <br>\nyou may need<br>\n split(\"\\\")[:-2]<br>\ninstead</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2262583,
      "author_name": "distiller",
      "author_url": "",
      "post_date": "05/17/2023 02:59:45",
      "content": "<p>how can i open it on an app instead of a jupyter kernel</p>",
      "votes": null,
      "replies": [
        {
          "id": 2262594,
          "author_name": "brimoor",
          "author_url": "",
          "post_date": "05/17/2023 03:32:08",
          "content": "<p>When you're working in Jupyter, you can launch the App in a separate tab rather than in the output cell like this:</p>\n<pre><code>\nsession = fo.launch_app(dataset, auto=)\nsession.open_tab()\n</code></pre>\n<p>Documentation: <a href=\"https://docs.voxel51.com/environments/index.html#opening-the-app-in-a-dedicated-tab\" target=\"_blank\">https://docs.voxel51.com/environments/index.html#opening-the-app-in-a-dedicated-tab</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2262643,
              "author_name": "distiller",
              "author_url": "",
              "post_date": "05/17/2023 04:26:56",
              "content": "<p>really helpful,thanks!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2226556": "Hi everyone, I put together a Python script to visualize and explore the dataset for this challenge. Check it out! 😄\n\nIf you find this helpful, consider giving FiftyOne a star on GitHub. It's free and open source!\nhttps://github.com/voxel51/fiftyone\n\n## Setup\n\n```\npip install fiftyone\n```\n\n## Load dataset\n\n```py\nimport glob\nimport fiftyone as fo\n\n# Download and unzip `image-matching-challenge-2023.zip` and put path here\ndataset_dir = \"/path/to/image-matching-challenge-2023\"\n\nsamples = []\nfor filepath in glob.glob(os.path.join(dataset_dir, \"**\"), recursive=True):\n    if filepath.endswith((\".jpg\", \".jpeg\", \".png\", \".JPG\")):\n        folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]\n        sample = fo.Sample(\n            filepath=filepath,\n            tags=[folders[0]],\n            type=folders[1],\n            location=folders[2],\n        )\n        samples.append(sample)\n\ndataset = fo.Dataset(\"image-matching-challenge-2023\", persistent=True)\ndataset.add_samples(samples)\ndataset.compute_metadata()\n```\n\n## Optional: add embeddings\n\n```py\nimport fiftyone.brain as fob\n\nfob.compute_visualization(\n    dataset,\n    model=\"clip-vit-base32-torch\",\n    brain_key=\"img_viz\",\n)\n```\n\n## Visualize in the App\n\n```py\nsession = fo.launch_app(dataset)\n```\n\n## Screenshots\n\nI could only upload images here, but I uploaded some screen recordings of me exploring the dataset on GitHub here: https://github.com/voxel51/fiftyone/issues/2906\n\n<img width=\"1132\" alt=\"Screen Shot 2023-04-18 at 11 12 30 PM\" src=\"https://user-images.githubusercontent.com/25985824/232957628-091a8e7a-3d45-4a55-8fee-30c366825cac.png\">\n\n<img width=\"2041\" alt=\"Screen Shot 2023-04-18 at 11 13 30 PM\" src=\"https://user-images.githubusercontent.com/25985824/232957624-f43b7c26-6f65-4a63-b291-d47a7aa07d1f.png\">",
    "2228134": "Great work! I put the star @brimoor",
    "2228691": "For detailed instructions on how to get FiftyOne installed, computing the embeddings and further data explorations, check out the full write up here: https://voxel51.com/blog/exploring-google-research-kaggle-image-matching-challenge-2023-dataset/",
    "2228697": "Thanks for the support! 🙌",
    "2262572": "folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]\nif you find something wrong on you PC,you may need folders = filepath[len(dataset_dir) + 1:].split(\"\\\\\")[:-2]instead",
    "2262576": "folders = filepath[len(dataset_dir) + 1:].split(\"/\")[:-2]\nif you find something wrong on you PC \nyou may need\n split(\"\\\\\")[:-2]\ninstead",
    "2262583": "how can i open it on an app instead of a jupyter kernel",
    "2262594": "When you're working in Jupyter, you can launch the App in a separate tab rather than in the output cell like this:\n\n```py\n# Launch the App in a dedicated browser tab\nsession = fo.launch_app(dataset, auto=False)\nsession.open_tab()\n```\n\nDocumentation: https://docs.voxel51.com/environments/index.html#opening-the-app-in-a-dedicated-tab",
    "2262643": "really helpful,thanks!"
  },
  "source": "meta"
}