{
  "id": 56670,
  "title": "Loading VGG16 Features",
  "url": "/competitions/avito-demand-prediction/discussion/56670",
  "author_name": "",
  "post_date": "2018-05-13T07:09:48.301859Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>Many kernels load the VGG16 features using the following code:</p>\n\n<pre><code>### Image features ###\ndef load_imfeatures(folder):\n    path = PurePath(folder)\n    features = sparse.load_npz(str(path / 'features.npz'))\n    return features\n\nftrain = load_imfeatures('../input/vgg16-train-features/')\nftest = load_imfeatures('../input/vgg16-test-features/')\n</code></pre>\n\n<p>When I run this on my Kaggle kernel, I get the error that <code>FileNotFoundError: [Errno 2] No such file or directory: '../input/vgg16-train-features/features.npz'</code> which to me makes sense since I have not imported any package containing the VGG16 features. How does this code work? </p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "328024",
      "postDate": "05/13/2018 07:09:48",
      "content": "<p>Hello,</p>\n\n<p>Many kernels load the VGG16 features using the following code:</p>\n\n<pre><code>### Image features ###\ndef load_imfeatures(folder):\n    path = PurePath(folder)\n    features = sparse.load_npz(str(path / 'features.npz'))\n    return features\n\nftrain = load_imfeatures('../input/vgg16-train-features/')\nftest = load_imfeatures('../input/vgg16-test-features/')\n</code></pre>\n\n<p>When I run this on my Kaggle kernel, I get the error that <code>FileNotFoundError: [Errno 2] No such file or directory: '../input/vgg16-train-features/features.npz'</code> which to me makes sense since I have not imported any package containing the VGG16 features. How does this code work? </p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hello,\n\nMany kernels load the VGG16 features using the following code:\n\n    ### Image features ###\n    def load_imfeatures(folder):\n        path = PurePath(folder)\n        features = sparse.load_npz(str(path / 'features.npz'))\n        return features\n    \n    ftrain = load_imfeatures('../input/vgg16-train-features/')\n    ftest = load_imfeatures('../input/vgg16-test-features/')\n\nWhen I run this on my Kaggle kernel, I get the error that `FileNotFoundError: [Errno 2] No such file or directory: '../input/vgg16-train-features/features.npz'` which to me makes sense since I have not imported any package containing the VGG16 features. How does this code work? \n\nThanks",
      "votes": null
    },
    {
      "id": "328035",
      "postDate": "05/13/2018 07:57:16",
      "content": "<p>You need to add one more Data from kernel's right panel.</p>",
      "rawMarkdown": "You need to add one more Data from kernel's right panel.",
      "votes": null
    },
    {
      "id": "328037",
      "postDate": "05/13/2018 07:58:25",
      "content": "<p>You need to add the output of <a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features/output\">Bruno G. do Amarals kernel</a> as data-source. </p>",
      "rawMarkdown": "You need to add the output of [Bruno G. do Amarals kernel][1] as data-source. \n\n\n  [1]: https://www.kaggle.com/bguberfain/vgg16-train-features/output",
      "votes": null
    },
    {
      "id": "328103",
      "postDate": "05/13/2018 10:50:03",
      "content": "<p>Thanks got it working. Then they use the following code:</p>\n\n<h1>Reduce image features</h1>\n\n<p>tsvd = TruncatedSVD(32)\nftsvd = tsvd.fit_transform(fboth)\ndel fboth\ngc.collect()</p>\n\n<h1>Merge image features into data</h1>\n\n<p>df_ftsvd = pd.DataFrame(ftsvd, index=df_both.index).add_prefix('im_tsvd_')\ndf_both = pd.concat([df_both, df_ftsvd], axis=1)\ndel df_ftsvd, ftsvd\ngc.collect()</p>\n\n<p>What does this do? I'm aware that it's some sort of dimensionality reduction from 512 (no features VGG16 expects) to 32, but I'm not sure how it works. </p>\n\n<p>Does the first bit just compress the 512 features into 32 and then the second bit looks for the presence of each of these features in each image in the dataset and appends a score of some sort to the dataframe based on this?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Thanks got it working. Then they use the following code:\n\n\n# Reduce image features\ntsvd = TruncatedSVD(32)\nftsvd = tsvd.fit_transform(fboth)\ndel fboth\ngc.collect()\n\n# Merge image features into data\ndf_ftsvd = pd.DataFrame(ftsvd, index=df_both.index).add_prefix('im_tsvd_')\ndf_both = pd.concat([df_both, df_ftsvd], axis=1)\ndel df_ftsvd, ftsvd\ngc.collect()\n\nWhat does this do? I'm aware that it's some sort of dimensionality reduction from 512 (no features VGG16 expects) to 32, but I'm not sure how it works. \n\nDoes the first bit just compress the 512 features into 32 and then the second bit looks for the presence of each of these features in each image in the dataset and appends a score of some sort to the dataframe based on this?\n\nThanks",
      "votes": null
    },
    {
      "id": "328284",
      "postDate": "05/13/2018 21:31:27",
      "content": "<p>It doesn't compress the 512 features into 32 binary features (presence vs. no presence)... it decomposes them into numeric vectors. <a href=\"https://en.wikipedia.org/wiki/Principal_component_analysis\">https://en.wikipedia.org/wiki/Principal_component_analysis</a></p>",
      "rawMarkdown": "It doesn't compress the 512 features into 32 binary features (presence vs. no presence)... it decomposes them into numeric vectors. https://en.wikipedia.org/wiki/Principal_component_analysis",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 328035,
      "author_name": "liujilong",
      "author_url": "",
      "post_date": "05/13/2018 07:57:16",
      "content": "<p>You need to add one more Data from kernel's right panel.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 328037,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "05/13/2018 07:58:25",
      "content": "<p>You need to add the output of <a href=\"https://www.kaggle.com/bguberfain/vgg16-train-features/output\">Bruno G. do Amarals kernel</a> as data-source. </p>",
      "votes": null,
      "replies": [
        {
          "id": 328103,
          "author_name": "derrington",
          "author_url": "",
          "post_date": "05/13/2018 10:50:03",
          "content": "<p>Thanks got it working. Then they use the following code:</p>\n\n<h1>Reduce image features</h1>\n\n<p>tsvd = TruncatedSVD(32)\nftsvd = tsvd.fit_transform(fboth)\ndel fboth\ngc.collect()</p>\n\n<h1>Merge image features into data</h1>\n\n<p>df_ftsvd = pd.DataFrame(ftsvd, index=df_both.index).add_prefix('im_tsvd_')\ndf_both = pd.concat([df_both, df_ftsvd], axis=1)\ndel df_ftsvd, ftsvd\ngc.collect()</p>\n\n<p>What does this do? I'm aware that it's some sort of dimensionality reduction from 512 (no features VGG16 expects) to 32, but I'm not sure how it works. </p>\n\n<p>Does the first bit just compress the 512 features into 32 and then the second bit looks for the presence of each of these features in each image in the dataset and appends a score of some sort to the dataframe based on this?</p>\n\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328284,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "05/13/2018 21:31:27",
          "content": "<p>It doesn't compress the 512 features into 32 binary features (presence vs. no presence)... it decomposes them into numeric vectors. <a href=\"https://en.wikipedia.org/wiki/Principal_component_analysis\">https://en.wikipedia.org/wiki/Principal_component_analysis</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "328024": "Hello,\n\nMany kernels load the VGG16 features using the following code:\n\n    ### Image features ###\n    def load_imfeatures(folder):\n        path = PurePath(folder)\n        features = sparse.load_npz(str(path / 'features.npz'))\n        return features\n    \n    ftrain = load_imfeatures('../input/vgg16-train-features/')\n    ftest = load_imfeatures('../input/vgg16-test-features/')\n\nWhen I run this on my Kaggle kernel, I get the error that `FileNotFoundError: [Errno 2] No such file or directory: '../input/vgg16-train-features/features.npz'` which to me makes sense since I have not imported any package containing the VGG16 features. How does this code work? \n\nThanks",
    "328035": "You need to add one more Data from kernel's right panel.",
    "328037": "You need to add the output of [Bruno G. do Amarals kernel][1] as data-source. \n\n\n  [1]: https://www.kaggle.com/bguberfain/vgg16-train-features/output",
    "328103": "Thanks got it working. Then they use the following code:\n\n\n# Reduce image features\ntsvd = TruncatedSVD(32)\nftsvd = tsvd.fit_transform(fboth)\ndel fboth\ngc.collect()\n\n# Merge image features into data\ndf_ftsvd = pd.DataFrame(ftsvd, index=df_both.index).add_prefix('im_tsvd_')\ndf_both = pd.concat([df_both, df_ftsvd], axis=1)\ndel df_ftsvd, ftsvd\ngc.collect()\n\nWhat does this do? I'm aware that it's some sort of dimensionality reduction from 512 (no features VGG16 expects) to 32, but I'm not sure how it works. \n\nDoes the first bit just compress the 512 features into 32 and then the second bit looks for the presence of each of these features in each image in the dataset and appends a score of some sort to the dataframe based on this?\n\nThanks",
    "328284": "It doesn't compress the 512 features into 32 binary features (presence vs. no presence)... it decomposes them into numeric vectors. https://en.wikipedia.org/wiki/Principal_component_analysis"
  },
  "source": "meta"
}