{
  "id": 165096,
  "title": "[Idea] Encoder-decoder with transformers (ct in, time series out) would be interesting",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/165096",
  "author_name": "",
  "post_date": "2020-07-08T14:52:34.967950300Z",
  "votes": 21,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Essentially just the title. </p>\n\n<p>This is just an idea, but essentially if you consider each CT scan as a single vector, then you can construct a permutation invariant set of vectors representing all the CT scans from the patient. You could encode the vector set with a transformer model (likely a small one, since the tensor representation from the CTs will take a lot of memory), and decode it with an auto-regressive transformer (similar to the original transformer approach), and at each time step you would generate a new FVC value. </p>\n\n<p>I think this could potentially offer a certain degree of interpretability since you can examine the attention heads, and retrace it to a certain slice (and using grad-CAM, you can even get which areas of the slice mattered the most for decoding the FVC value at a specific time step).</p>",
  "messages": [
    {
      "id": "920375",
      "postDate": "07/08/2020 14:52:34",
      "content": "<p>Essentially just the title. </p>\n\n<p>This is just an idea, but essentially if you consider each CT scan as a single vector, then you can construct a permutation invariant set of vectors representing all the CT scans from the patient. You could encode the vector set with a transformer model (likely a small one, since the tensor representation from the CTs will take a lot of memory), and decode it with an auto-regressive transformer (similar to the original transformer approach), and at each time step you would generate a new FVC value. </p>\n\n<p>I think this could potentially offer a certain degree of interpretability since you can examine the attention heads, and retrace it to a certain slice (and using grad-CAM, you can even get which areas of the slice mattered the most for decoding the FVC value at a specific time step).</p>",
      "rawMarkdown": "Essentially just the title. \n\nThis is just an idea, but essentially if you consider each CT scan as a single vector, then you can construct a permutation invariant set of vectors representing all the CT scans from the patient. You could encode the vector set with a transformer model (likely a small one, since the tensor representation from the CTs will take a lot of memory), and decode it with an auto-regressive transformer (similar to the original transformer approach), and at each time step you would generate a new FVC value. \n\nI think this could potentially offer a certain degree of interpretability since you can examine the attention heads, and retrace it to a certain slice (and using grad-CAM, you can even get which areas of the slice mattered the most for decoding the FVC value at a specific time step).",
      "votes": null
    },
    {
      "id": "925388",
      "postDate": "07/12/2020 04:07:20",
      "content": "<p>it is</p>",
      "rawMarkdown": "it is",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 925388,
      "author_name": "cboychinedu",
      "author_url": "",
      "post_date": "07/12/2020 04:07:20",
      "content": "<p>it is</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "920375": "Essentially just the title. \n\nThis is just an idea, but essentially if you consider each CT scan as a single vector, then you can construct a permutation invariant set of vectors representing all the CT scans from the patient. You could encode the vector set with a transformer model (likely a small one, since the tensor representation from the CTs will take a lot of memory), and decode it with an auto-regressive transformer (similar to the original transformer approach), and at each time step you would generate a new FVC value. \n\nI think this could potentially offer a certain degree of interpretability since you can examine the attention heads, and retrace it to a certain slice (and using grad-CAM, you can even get which areas of the slice mattered the most for decoding the FVC value at a specific time step).",
    "925388": "it is"
  },
  "source": "meta"
}