{
  "id": 350167,
  "title": "MSCI Multiome: handling the full dataset with pytorch",
  "url": "/competitions/open-problems-multimodal/discussion/350167",
  "author_name": "",
  "post_date": "2022-09-04T14:10:20.232338300Z",
  "votes": 24,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>As a followup to <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349031\" target=\"_blank\">this topic</a>, I have tried to use pytorch support for sparse tensors to train a Neural Network on the full multiome dataset.</p>\n<p>It turns out it is actually realistic to load the full dataset in GPU memory; leaving 4GB free to train a model. It is also possible to keep the dataset in RAM (leaving 1GB of free RAM) and keep the 16GB GPU memory available.</p>\n<p>For now, given that we only have ~100K training examples, I do not think we will need very large models, therefore I have taken the option of putting the full dataset in GPU memory, which is more efficient.</p>\n<p>Since I think the functions I wrote can be useful for other kagglers wanting to apply DNN to Multiome, I have made public two \"proof-of-concept\" notebooks using a simple MLP applied on the raw Multiome inputs (it might be better to apply dimensionality reduction first, but I think it is interesting to let the neural network see the full data).</p>\n<p>Training notebook:<br>\n<a href=\"https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors\" target=\"_blank\">https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors</a></p>\n<p>Submission notebook:<br>\n<a href=\"https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-submission\" target=\"_blank\">https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-submission</a></p>\n<p>At this point, the LB result of these notebooks is a bit below the much simpler PCA+Ridge Regression method. But this is a much more flexible approach, so I hope it can lead to significant improvements after a bit of work.</p>",
  "messages": [
    {
      "id": "1926043",
      "postDate": "09/04/2022 14:10:20",
      "content": "<p>Hi,</p>\n<p>As a followup to <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349031\" target=\"_blank\">this topic</a>, I have tried to use pytorch support for sparse tensors to train a Neural Network on the full multiome dataset.</p>\n<p>It turns out it is actually realistic to load the full dataset in GPU memory; leaving 4GB free to train a model. It is also possible to keep the dataset in RAM (leaving 1GB of free RAM) and keep the 16GB GPU memory available.</p>\n<p>For now, given that we only have ~100K training examples, I do not think we will need very large models, therefore I have taken the option of putting the full dataset in GPU memory, which is more efficient.</p>\n<p>Since I think the functions I wrote can be useful for other kagglers wanting to apply DNN to Multiome, I have made public two \"proof-of-concept\" notebooks using a simple MLP applied on the raw Multiome inputs (it might be better to apply dimensionality reduction first, but I think it is interesting to let the neural network see the full data).</p>\n<p>Training notebook:<br>\n<a href=\"https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors\" target=\"_blank\">https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors</a></p>\n<p>Submission notebook:<br>\n<a href=\"https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-submission\" target=\"_blank\">https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-submission</a></p>\n<p>At this point, the LB result of these notebooks is a bit below the much simpler PCA+Ridge Regression method. But this is a much more flexible approach, so I hope it can lead to significant improvements after a bit of work.</p>",
      "rawMarkdown": "Hi,\n\nAs a followup to [this topic](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349031), I have tried to use pytorch support for sparse tensors to train a Neural Network on the full multiome dataset.\n\nIt turns out it is actually realistic to load the full dataset in GPU memory; leaving 4GB free to train a model. It is also possible to keep the dataset in RAM (leaving 1GB of free RAM) and keep the 16GB GPU memory available.\n\nFor now, given that we only have ~100K training examples, I do not think we will need very large models, therefore I have taken the option of putting the full dataset in GPU memory, which is more efficient.\n\nSince I think the functions I wrote can be useful for other kagglers wanting to apply DNN to Multiome, I have made public two \"proof-of-concept\" notebooks using a simple MLP applied on the raw Multiome inputs (it might be better to apply dimensionality reduction first, but I think it is interesting to let the neural network see the full data).\n\nTraining notebook:\nhttps://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors\n\nSubmission notebook:\nhttps://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-submission\n\nAt this point, the LB result of these notebooks is a bit below the much simpler PCA+Ridge Regression method. But this is a much more flexible approach, so I hope it can lead to significant improvements after a bit of work.",
      "votes": null
    },
    {
      "id": "1926287",
      "postDate": "09/04/2022 17:56:53",
      "content": "<p>Thank you very much for sharing ! </p>",
      "rawMarkdown": "Thank you very much for sharing !",
      "votes": null
    },
    {
      "id": "2015566",
      "postDate": "11/03/2022 11:36:41",
      "content": "<p>Thank you very much for sharing !</p>",
      "rawMarkdown": "Thank you very much for sharing !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1926287,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/04/2022 17:56:53",
      "content": "<p>Thank you very much for sharing ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2015566,
      "author_name": "lcbupt",
      "author_url": "",
      "post_date": "11/03/2022 11:36:41",
      "content": "<p>Thank you very much for sharing !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1926043": "Hi,\n\nAs a followup to [this topic](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/349031), I have tried to use pytorch support for sparse tensors to train a Neural Network on the full multiome dataset.\n\nIt turns out it is actually realistic to load the full dataset in GPU memory; leaving 4GB free to train a model. It is also possible to keep the dataset in RAM (leaving 1GB of free RAM) and keep the 16GB GPU memory available.\n\nFor now, given that we only have ~100K training examples, I do not think we will need very large models, therefore I have taken the option of putting the full dataset in GPU memory, which is more efficient.\n\nSince I think the functions I wrote can be useful for other kagglers wanting to apply DNN to Multiome, I have made public two \"proof-of-concept\" notebooks using a simple MLP applied on the raw Multiome inputs (it might be better to apply dimensionality reduction first, but I think it is interesting to let the neural network see the full data).\n\nTraining notebook:\nhttps://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors\n\nSubmission notebook:\nhttps://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-submission\n\nAt this point, the LB result of these notebooks is a bit below the much simpler PCA+Ridge Regression method. But this is a much more flexible approach, so I hope it can lead to significant improvements after a bit of work.",
    "1926287": "Thank you very much for sharing !",
    "2015566": "Thank you very much for sharing !"
  },
  "source": "meta"
}