{
  "id": 296464,
  "title": "NN Easily Fooled by Strange Poses. Not Madonna. But Inception, Yolo and Coco.",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/296464",
  "author_name": "",
  "post_date": "2021-12-21T18:37:31.680584800Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Strike (with) a Pose: Neural Networks Are Easily Fooled by Strange Poses of Familiar Objects</h1>\n<p>Authors: Michael A. Alcorn, Qi Li, Zhitao Gong, Chengfei Wang, Long Mai, Wei-Shinn Ku, Anh Nguyen</p>\n<p>arXiv:1811.11553v3 [cs.CV] </p>\n<p>\"Despite excellent performance on stationary test sets, deep neural networks (DNNs) can fail to generalize to out-of-distribution (OoD) inputs, including natural, non-adversarial ones, which are common in real-world settings.\"</p>\n<p>\"In that paper, the authors presented a framework for discovering DNN failures that harnesses 3D renderers and 3D models. That is, they estimated the parameters of a 3D renderer that cause a target DNN to misbehave in response to the rendered image.</p>\n<p>\"Using their framework and a self-assembled dataset of 3D objects, they investigated the vulnerability of DNNs to OoD poses of well-known objects in ImageNet. For objects that are readily recognized by DNNs in their canonical poses, DNNs incorrectly classify 97% of their pose space.\"</p>\n<p>\"In addition, DNNs are highly sensitive to slight pose perturbations. Importantly, adversarial poses transfer across models and datasets. The authors found that 99.9% and 99.4% of the poses misclassified by Inception-v3 also transfer to the AlexNet and ResNet-50 image classifiers trained on the same ImageNet dataset, respectively, and 75.5% transfer to the YOLOv3 object detector trained on MS COCO.</p>\n<p>Poster at the 2019 Conference on Computer Vision and Pattern Recognition<br>\n<a href=\"https://arxiv.org/abs/1811.11553\" target=\"_blank\">https://arxiv.org/abs/1811.11553</a></p>\n<h1>Strike a pose to control COTS (Crown-of-thorns-Starfish)!</h1>",
  "messages": [
    {
      "id": "1625396",
      "postDate": "12/21/2021 18:37:31",
      "content": "<h1>Strike (with) a Pose: Neural Networks Are Easily Fooled by Strange Poses of Familiar Objects</h1>\n<p>Authors: Michael A. Alcorn, Qi Li, Zhitao Gong, Chengfei Wang, Long Mai, Wei-Shinn Ku, Anh Nguyen</p>\n<p>arXiv:1811.11553v3 [cs.CV] </p>\n<p>\"Despite excellent performance on stationary test sets, deep neural networks (DNNs) can fail to generalize to out-of-distribution (OoD) inputs, including natural, non-adversarial ones, which are common in real-world settings.\"</p>\n<p>\"In that paper, the authors presented a framework for discovering DNN failures that harnesses 3D renderers and 3D models. That is, they estimated the parameters of a 3D renderer that cause a target DNN to misbehave in response to the rendered image.</p>\n<p>\"Using their framework and a self-assembled dataset of 3D objects, they investigated the vulnerability of DNNs to OoD poses of well-known objects in ImageNet. For objects that are readily recognized by DNNs in their canonical poses, DNNs incorrectly classify 97% of their pose space.\"</p>\n<p>\"In addition, DNNs are highly sensitive to slight pose perturbations. Importantly, adversarial poses transfer across models and datasets. The authors found that 99.9% and 99.4% of the poses misclassified by Inception-v3 also transfer to the AlexNet and ResNet-50 image classifiers trained on the same ImageNet dataset, respectively, and 75.5% transfer to the YOLOv3 object detector trained on MS COCO.</p>\n<p>Poster at the 2019 Conference on Computer Vision and Pattern Recognition<br>\n<a href=\"https://arxiv.org/abs/1811.11553\" target=\"_blank\">https://arxiv.org/abs/1811.11553</a></p>\n<h1>Strike a pose to control COTS (Crown-of-thorns-Starfish)!</h1>",
      "rawMarkdown": "#Strike (with) a Pose: Neural Networks Are Easily Fooled by Strange Poses of Familiar Objects\n\nAuthors: Michael A. Alcorn, Qi Li, Zhitao Gong, Chengfei Wang, Long Mai, Wei-Shinn Ku, Anh Nguyen\n\n arXiv:1811.11553v3 [cs.CV] \n\n\"Despite excellent performance on stationary test sets, deep neural networks (DNNs) can fail to generalize to out-of-distribution (OoD) inputs, including natural, non-adversarial ones, which are common in real-world settings.\"\n\n\"In that paper, the authors presented a framework for discovering DNN failures that harnesses 3D renderers and 3D models. That is, they estimated the parameters of a 3D renderer that cause a target DNN to misbehave in response to the rendered image.\n\n\"Using their framework and a self-assembled dataset of 3D objects, they investigated the vulnerability of DNNs to OoD poses of well-known objects in ImageNet. For objects that are readily recognized by DNNs in their canonical poses, DNNs incorrectly classify 97% of their pose space.\"\n\n\"In addition, DNNs are highly sensitive to slight pose perturbations. Importantly, adversarial poses transfer across models and datasets. The authors found that 99.9% and 99.4% of the poses misclassified by Inception-v3 also transfer to the AlexNet and ResNet-50 image classifiers trained on the same ImageNet dataset, respectively, and 75.5% transfer to the YOLOv3 object detector trained on MS COCO.\n\n\nPoster at the 2019 Conference on Computer Vision and Pattern Recognition\nhttps://arxiv.org/abs/1811.11553\n\n#Strike a pose to control COTS (Crown-of-thorns-Starfish)!",
      "votes": null
    },
    {
      "id": "1625940",
      "postDate": "12/22/2021 10:01:06",
      "content": "<p>This is a link to a paper, from DeepMind, that attempts to address this problem. <br>\n<a href=\"https://arxiv.org/abs/2106.05886\" target=\"_blank\">https://arxiv.org/abs/2106.05886</a></p>\n<p>The idea is to modify the input to CNN layers to make the model robust to translations and rotations. In essence, given two versions of the same image, the model should produce nearly identical vector representations of each. </p>\n<p>I believe this robustness would be helpful when building vision systems for drones and submersibles - they are able to move in three dimensions and may therefore view their subjects from unusual angles.</p>",
      "rawMarkdown": "This is a link to a paper, from DeepMind, that attempts to address this problem. \nhttps://arxiv.org/abs/2106.05886\n\nThe idea is to modify the input to CNN layers to make the model robust to translations and rotations. In essence, given two versions of the same image, the model should produce nearly identical vector representations of each. \n\nI believe this robustness would be helpful when building vision systems for drones and submersibles - they are able to move in three dimensions and may therefore view their subjects from unusual angles.",
      "votes": null
    },
    {
      "id": "1626123",
      "postDate": "12/22/2021 13:29:00",
      "content": "<p>Hi Marsh,</p>\n<p>Thank you with the link (Group Equivariant Subsampling in CNNs) and the support too. </p>",
      "rawMarkdown": "Hi Marsh,\n\nThank you with the link (Group Equivariant Subsampling in CNNs) and the support too.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1625940,
      "author_name": "vbookshelf",
      "author_url": "",
      "post_date": "12/22/2021 10:01:06",
      "content": "<p>This is a link to a paper, from DeepMind, that attempts to address this problem. <br>\n<a href=\"https://arxiv.org/abs/2106.05886\" target=\"_blank\">https://arxiv.org/abs/2106.05886</a></p>\n<p>The idea is to modify the input to CNN layers to make the model robust to translations and rotations. In essence, given two versions of the same image, the model should produce nearly identical vector representations of each. </p>\n<p>I believe this robustness would be helpful when building vision systems for drones and submersibles - they are able to move in three dimensions and may therefore view their subjects from unusual angles.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1626123,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "12/22/2021 13:29:00",
          "content": "<p>Hi Marsh,</p>\n<p>Thank you with the link (Group Equivariant Subsampling in CNNs) and the support too. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1625396": "#Strike (with) a Pose: Neural Networks Are Easily Fooled by Strange Poses of Familiar Objects\n\nAuthors: Michael A. Alcorn, Qi Li, Zhitao Gong, Chengfei Wang, Long Mai, Wei-Shinn Ku, Anh Nguyen\n\n arXiv:1811.11553v3 [cs.CV] \n\n\"Despite excellent performance on stationary test sets, deep neural networks (DNNs) can fail to generalize to out-of-distribution (OoD) inputs, including natural, non-adversarial ones, which are common in real-world settings.\"\n\n\"In that paper, the authors presented a framework for discovering DNN failures that harnesses 3D renderers and 3D models. That is, they estimated the parameters of a 3D renderer that cause a target DNN to misbehave in response to the rendered image.\n\n\"Using their framework and a self-assembled dataset of 3D objects, they investigated the vulnerability of DNNs to OoD poses of well-known objects in ImageNet. For objects that are readily recognized by DNNs in their canonical poses, DNNs incorrectly classify 97% of their pose space.\"\n\n\"In addition, DNNs are highly sensitive to slight pose perturbations. Importantly, adversarial poses transfer across models and datasets. The authors found that 99.9% and 99.4% of the poses misclassified by Inception-v3 also transfer to the AlexNet and ResNet-50 image classifiers trained on the same ImageNet dataset, respectively, and 75.5% transfer to the YOLOv3 object detector trained on MS COCO.\n\n\nPoster at the 2019 Conference on Computer Vision and Pattern Recognition\nhttps://arxiv.org/abs/1811.11553\n\n#Strike a pose to control COTS (Crown-of-thorns-Starfish)!",
    "1625940": "This is a link to a paper, from DeepMind, that attempts to address this problem. \nhttps://arxiv.org/abs/2106.05886\n\nThe idea is to modify the input to CNN layers to make the model robust to translations and rotations. In essence, given two versions of the same image, the model should produce nearly identical vector representations of each. \n\nI believe this robustness would be helpful when building vision systems for drones and submersibles - they are able to move in three dimensions and may therefore view their subjects from unusual angles.",
    "1626123": "Hi Marsh,\n\nThank you with the link (Group Equivariant Subsampling in CNNs) and the support too."
  },
  "source": "meta"
}