{
  "id": 398430,
  "title": "Sharing Masked Autoencoder Idea",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/398430",
  "author_name": "",
  "post_date": "2023-03-30T03:22:40.351462400Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all—Ben here. I wanted to share something I've been working on towards the Ink Detection challenge that I thought could be really cool if other people built off it :) I played with it for a day or so and wasn't able to get the loss to go down like I wanted, but I think the concept is potentially promising!</p>\n<p>The basic idea is the \"masked auto-encoder\" approach to Vision Transformers introduced in the paper \"Masked Autoencoders are Scalable Vision Learners\" (<a href=\"https://arxiv.org/abs/2111.06377)\" target=\"_blank\">https://arxiv.org/abs/2111.06377)</a>. It's a self-supervised pretraining approach that works by having the model predict masked patches of an image (or in this case, a chunk of a voxel). My thought was by doing self-supervised pretraining on voxel slices, this model would learn something about the fine-grained representations of the tomography scans, which would then transfer to the downstream task. </p>\n<p>I've had some trouble getting it to converge on the three provided fragments, but I think it's worth a shot if someone can build on it! Even if that approach doesn't work out, you might be interested in my IterableDataset that  can \"stream\" sub-patches from the image files and possibly avoid some memory issues, if you want to only use part of a \"stack\" at a time.</p>\n<p>The Colab with all of that code is right here: <a href=\"https://colab.research.google.com/drive/1bbr1iKceb6Op_BEoiNvY9nl9U_akPPNF?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1bbr1iKceb6Op_BEoiNvY9nl9U_akPPNF?usp=sharing</a></p>",
  "messages": [
    {
      "id": "2202407",
      "postDate": "03/30/2023 03:22:40",
      "content": "<p>Hi all—Ben here. I wanted to share something I've been working on towards the Ink Detection challenge that I thought could be really cool if other people built off it :) I played with it for a day or so and wasn't able to get the loss to go down like I wanted, but I think the concept is potentially promising!</p>\n<p>The basic idea is the \"masked auto-encoder\" approach to Vision Transformers introduced in the paper \"Masked Autoencoders are Scalable Vision Learners\" (<a href=\"https://arxiv.org/abs/2111.06377)\" target=\"_blank\">https://arxiv.org/abs/2111.06377)</a>. It's a self-supervised pretraining approach that works by having the model predict masked patches of an image (or in this case, a chunk of a voxel). My thought was by doing self-supervised pretraining on voxel slices, this model would learn something about the fine-grained representations of the tomography scans, which would then transfer to the downstream task. </p>\n<p>I've had some trouble getting it to converge on the three provided fragments, but I think it's worth a shot if someone can build on it! Even if that approach doesn't work out, you might be interested in my IterableDataset that  can \"stream\" sub-patches from the image files and possibly avoid some memory issues, if you want to only use part of a \"stack\" at a time.</p>\n<p>The Colab with all of that code is right here: <a href=\"https://colab.research.google.com/drive/1bbr1iKceb6Op_BEoiNvY9nl9U_akPPNF?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1bbr1iKceb6Op_BEoiNvY9nl9U_akPPNF?usp=sharing</a></p>",
      "rawMarkdown": "Hi all—Ben here. I wanted to share something I've been working on towards the Ink Detection challenge that I thought could be really cool if other people built off it :) I played with it for a day or so and wasn't able to get the loss to go down like I wanted, but I think the concept is potentially promising!\n\nThe basic idea is the \"masked auto-encoder\" approach to Vision Transformers introduced in the paper \"Masked Autoencoders are Scalable Vision Learners\" (https://arxiv.org/abs/2111.06377). It's a self-supervised pretraining approach that works by having the model predict masked patches of an image (or in this case, a chunk of a voxel). My thought was by doing self-supervised pretraining on voxel slices, this model would learn something about the fine-grained representations of the tomography scans, which would then transfer to the downstream task. \n\nI've had some trouble getting it to converge on the three provided fragments, but I think it's worth a shot if someone can build on it! Even if that approach doesn't work out, you might be interested in my IterableDataset that  can \"stream\" sub-patches from the image files and possibly avoid some memory issues, if you want to only use part of a \"stack\" at a time.\n\nThe Colab with all of that code is right here: https://colab.research.google.com/drive/1bbr1iKceb6Op_BEoiNvY9nl9U_akPPNF?usp=sharing",
      "votes": null
    },
    {
      "id": "2203249",
      "postDate": "03/30/2023 16:38:32",
      "content": "<p>Honestly, I don't think MAE will work here. If a masked image has a patch containing an animal tail, then it is very likely the image has some other animal body parts in it (correlation), which narrow down the probability distribution from the one corresponds to a random ImageNet image (large reduce in information entropy). I think MAE relies on the long-range correlation of image pixels and sparse density of information, both of which do not exist in the Vesuvius dataset.</p>\n<p>Of course I could be wrong, and welcome to discussions.</p>",
      "rawMarkdown": "Honestly, I don't think MAE will work here. If a masked image has a patch containing an animal tail, then it is very likely the image has some other animal body parts in it (correlation), which narrow down the probability distribution from the one corresponds to a random ImageNet image (large reduce in information entropy). I think MAE relies on the long-range correlation of image pixels and sparse density of information, both of which do not exist in the Vesuvius dataset.\n\nOf course I could be wrong, and welcome to discussions.",
      "votes": null
    },
    {
      "id": "2213436",
      "postDate": "04/07/2023 15:17:00",
      "content": "<p>Hi Ben,</p>\n<p>Thanks for sharing your Colab link and your autoencoder idea.</p>",
      "rawMarkdown": "Hi Ben,\n\nThanks for sharing your Colab link and your autoencoder idea.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2203249,
      "author_name": "junxhuang",
      "author_url": "",
      "post_date": "03/30/2023 16:38:32",
      "content": "<p>Honestly, I don't think MAE will work here. If a masked image has a patch containing an animal tail, then it is very likely the image has some other animal body parts in it (correlation), which narrow down the probability distribution from the one corresponds to a random ImageNet image (large reduce in information entropy). I think MAE relies on the long-range correlation of image pixels and sparse density of information, both of which do not exist in the Vesuvius dataset.</p>\n<p>Of course I could be wrong, and welcome to discussions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2213436,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "04/07/2023 15:17:00",
      "content": "<p>Hi Ben,</p>\n<p>Thanks for sharing your Colab link and your autoencoder idea.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2202407": "Hi all—Ben here. I wanted to share something I've been working on towards the Ink Detection challenge that I thought could be really cool if other people built off it :) I played with it for a day or so and wasn't able to get the loss to go down like I wanted, but I think the concept is potentially promising!\n\nThe basic idea is the \"masked auto-encoder\" approach to Vision Transformers introduced in the paper \"Masked Autoencoders are Scalable Vision Learners\" (https://arxiv.org/abs/2111.06377). It's a self-supervised pretraining approach that works by having the model predict masked patches of an image (or in this case, a chunk of a voxel). My thought was by doing self-supervised pretraining on voxel slices, this model would learn something about the fine-grained representations of the tomography scans, which would then transfer to the downstream task. \n\nI've had some trouble getting it to converge on the three provided fragments, but I think it's worth a shot if someone can build on it! Even if that approach doesn't work out, you might be interested in my IterableDataset that  can \"stream\" sub-patches from the image files and possibly avoid some memory issues, if you want to only use part of a \"stack\" at a time.\n\nThe Colab with all of that code is right here: https://colab.research.google.com/drive/1bbr1iKceb6Op_BEoiNvY9nl9U_akPPNF?usp=sharing",
    "2203249": "Honestly, I don't think MAE will work here. If a masked image has a patch containing an animal tail, then it is very likely the image has some other animal body parts in it (correlation), which narrow down the probability distribution from the one corresponds to a random ImageNet image (large reduce in information entropy). I think MAE relies on the long-range correlation of image pixels and sparse density of information, both of which do not exist in the Vesuvius dataset.\n\nOf course I could be wrong, and welcome to discussions.",
    "2213436": "Hi Ben,\n\nThanks for sharing your Colab link and your autoencoder idea."
  },
  "source": "meta"
}