{
  "id": 349841,
  "title": "How to start？",
  "url": "/competitions/open-problems-multimodal/discussion/349841",
  "author_name": "",
  "post_date": "2022-09-03T02:49:25.552287500Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone, I'm a beginner in this field. I'm preparing for the competition recently, but I don't know how to start. Anyone willing to give a tutorial? Does it need to use deep learning or machine learning?</p>",
  "messages": [
    {
      "id": "1924379",
      "postDate": "09/03/2022 02:49:25",
      "content": "<p>Hello everyone, I'm a beginner in this field. I'm preparing for the competition recently, but I don't know how to start. Anyone willing to give a tutorial? Does it need to use deep learning or machine learning?</p>",
      "rawMarkdown": "Hello everyone, I'm a beginner in this field. I'm preparing for the competition recently, but I don't know how to start. Anyone willing to give a tutorial? Does it need to use deep learning or machine learning?",
      "votes": null
    },
    {
      "id": "1924460",
      "postDate": "09/03/2022 04:11:47",
      "content": "<p>Hi! <a href=\"https://www.kaggle.com/cuidongdong\" target=\"_blank\">@cuidongdong</a> !<br>\nI myself have no knowledge of cell analysis and bioinformatics but you can still compete and make good predictions. This is also my first attempt on a fully-fledged high-level competition.<br>\nI will attempt to explain what I have gathered from discussions and notebooks, without bringing in any domain knowledge of bioinformatics. I have deduced following points (any reader can correct me if I am wrong on any point):</p>\n<ul>\n<li>The competition can be seen as split in two parts: Citeseq and Multiome. We have to prepare two models to make predictions on Citeseq and Multiome. So, you can treat these datasets differently.</li>\n<li>You can look at this competition as a multioutput regression problem where based on the data we have to predict multiple targets. The evaluation metric is <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/overview/evaluation\" target=\"_blank\">Pearson's Correlation Coefficient</a>.</li>\n<li>Assuming by machine learning you mean simple regression models like Ridge, Lasso etc. and by deep learning Neural Networks, you can use any approach.</li>\n</ul>\n<p>Currently, I have only looked at Citeseq data and I think following points are going to be important:</p>\n<ul>\n<li>The cross-validation strategy is going to be important as the test data has an unknown donor and a different day.</li>\n<li>Dimensionality Reduction and Feature Analysis are also going to be important as there are many features to select from.</li>\n<li>While I have not brought up any points from domain knowledge, it is going to play an important role I believe as it will be helpful to interpret data and extract more meaningful information and features. So, keep an eye on discussions as people are bringing in many important points.</li>\n</ul>\n<p>Some resources that can be helpful to start:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346888\" target=\"_blank\">Some Domain and Competition Knowledge</a></li>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\" target=\"_blank\">Citeseq Quickstart</a></li>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/msci-multiome-quickstart\" target=\"_blank\">Multiome Quickstart</a></li>\n<li><a href=\"https://github.com/openproblems-bio/neurips_2022_saturn_notebooks\" target=\"_blank\">NeurIPS Notebooks Github</a> as it contains some very good starter notebooks.</li>\n</ul>\n<p>Good luck!</p>",
      "rawMarkdown": "Hi! @cuidongdong !\nI myself have no knowledge of cell analysis and bioinformatics but you can still compete and make good predictions. This is also my first attempt on a fully-fledged high-level competition.\nI will attempt to explain what I have gathered from discussions and notebooks, without bringing in any domain knowledge of bioinformatics. I have deduced following points (any reader can correct me if I am wrong on any point):\n- The competition can be seen as split in two parts: Citeseq and Multiome. We have to prepare two models to make predictions on Citeseq and Multiome. So, you can treat these datasets differently.\n- You can look at this competition as a multioutput regression problem where based on the data we have to predict multiple targets. The evaluation metric is [Pearson's Correlation Coefficient](https://www.kaggle.com/competitions/open-problems-multimodal/overview/evaluation).\n- Assuming by machine learning you mean simple regression models like Ridge, Lasso etc. and by deep learning Neural Networks, you can use any approach.\n\nCurrently, I have only looked at Citeseq data and I think following points are going to be important:\n- The cross-validation strategy is going to be important as the test data has an unknown donor and a different day.\n- Dimensionality Reduction and Feature Analysis are also going to be important as there are many features to select from.\n- While I have not brought up any points from domain knowledge, it is going to play an important role I believe as it will be helpful to interpret data and extract more meaningful information and features. So, keep an eye on discussions as people are bringing in many important points.\n\nSome resources that can be helpful to start:\n- [Some Domain and Competition Knowledge](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346888)\n- [Citeseq Quickstart](https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart)\n- [Multiome Quickstart](https://www.kaggle.com/code/ambrosm/msci-multiome-quickstart)\n- [NeurIPS Notebooks Github](https://github.com/openproblems-bio/neurips_2022_saturn_notebooks) as it contains some very good starter notebooks.\n\nGood luck!",
      "votes": null
    },
    {
      "id": "1924534",
      "postDate": "09/03/2022 06:30:42",
      "content": "<p><a href=\"https://youtu.be/iTWdjtaSf9g\" target=\"_blank\">https://youtu.be/iTWdjtaSf9g</a><br>\nAndrew Lukyanenko (Kaggle Grandmaster) \"Getting started with Kaggle\"</p>",
      "rawMarkdown": "https://youtu.be/iTWdjtaSf9g\nAndrew Lukyanenko (Kaggle Grandmaster) \"Getting started with Kaggle\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1924460,
      "author_name": "sadmadlad",
      "author_url": "",
      "post_date": "09/03/2022 04:11:47",
      "content": "<p>Hi! <a href=\"https://www.kaggle.com/cuidongdong\" target=\"_blank\">@cuidongdong</a> !<br>\nI myself have no knowledge of cell analysis and bioinformatics but you can still compete and make good predictions. This is also my first attempt on a fully-fledged high-level competition.<br>\nI will attempt to explain what I have gathered from discussions and notebooks, without bringing in any domain knowledge of bioinformatics. I have deduced following points (any reader can correct me if I am wrong on any point):</p>\n<ul>\n<li>The competition can be seen as split in two parts: Citeseq and Multiome. We have to prepare two models to make predictions on Citeseq and Multiome. So, you can treat these datasets differently.</li>\n<li>You can look at this competition as a multioutput regression problem where based on the data we have to predict multiple targets. The evaluation metric is <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/overview/evaluation\" target=\"_blank\">Pearson's Correlation Coefficient</a>.</li>\n<li>Assuming by machine learning you mean simple regression models like Ridge, Lasso etc. and by deep learning Neural Networks, you can use any approach.</li>\n</ul>\n<p>Currently, I have only looked at Citeseq data and I think following points are going to be important:</p>\n<ul>\n<li>The cross-validation strategy is going to be important as the test data has an unknown donor and a different day.</li>\n<li>Dimensionality Reduction and Feature Analysis are also going to be important as there are many features to select from.</li>\n<li>While I have not brought up any points from domain knowledge, it is going to play an important role I believe as it will be helpful to interpret data and extract more meaningful information and features. So, keep an eye on discussions as people are bringing in many important points.</li>\n</ul>\n<p>Some resources that can be helpful to start:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346888\" target=\"_blank\">Some Domain and Competition Knowledge</a></li>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\" target=\"_blank\">Citeseq Quickstart</a></li>\n<li><a href=\"https://www.kaggle.com/code/ambrosm/msci-multiome-quickstart\" target=\"_blank\">Multiome Quickstart</a></li>\n<li><a href=\"https://github.com/openproblems-bio/neurips_2022_saturn_notebooks\" target=\"_blank\">NeurIPS Notebooks Github</a> as it contains some very good starter notebooks.</li>\n</ul>\n<p>Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1924534,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/03/2022 06:30:42",
      "content": "<p><a href=\"https://youtu.be/iTWdjtaSf9g\" target=\"_blank\">https://youtu.be/iTWdjtaSf9g</a><br>\nAndrew Lukyanenko (Kaggle Grandmaster) \"Getting started with Kaggle\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1924379": "Hello everyone, I'm a beginner in this field. I'm preparing for the competition recently, but I don't know how to start. Anyone willing to give a tutorial? Does it need to use deep learning or machine learning?",
    "1924460": "Hi! @cuidongdong !\nI myself have no knowledge of cell analysis and bioinformatics but you can still compete and make good predictions. This is also my first attempt on a fully-fledged high-level competition.\nI will attempt to explain what I have gathered from discussions and notebooks, without bringing in any domain knowledge of bioinformatics. I have deduced following points (any reader can correct me if I am wrong on any point):\n- The competition can be seen as split in two parts: Citeseq and Multiome. We have to prepare two models to make predictions on Citeseq and Multiome. So, you can treat these datasets differently.\n- You can look at this competition as a multioutput regression problem where based on the data we have to predict multiple targets. The evaluation metric is [Pearson's Correlation Coefficient](https://www.kaggle.com/competitions/open-problems-multimodal/overview/evaluation).\n- Assuming by machine learning you mean simple regression models like Ridge, Lasso etc. and by deep learning Neural Networks, you can use any approach.\n\nCurrently, I have only looked at Citeseq data and I think following points are going to be important:\n- The cross-validation strategy is going to be important as the test data has an unknown donor and a different day.\n- Dimensionality Reduction and Feature Analysis are also going to be important as there are many features to select from.\n- While I have not brought up any points from domain knowledge, it is going to play an important role I believe as it will be helpful to interpret data and extract more meaningful information and features. So, keep an eye on discussions as people are bringing in many important points.\n\nSome resources that can be helpful to start:\n- [Some Domain and Competition Knowledge](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346888)\n- [Citeseq Quickstart](https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart)\n- [Multiome Quickstart](https://www.kaggle.com/code/ambrosm/msci-multiome-quickstart)\n- [NeurIPS Notebooks Github](https://github.com/openproblems-bio/neurips_2022_saturn_notebooks) as it contains some very good starter notebooks.\n\nGood luck!",
    "1924534": "https://youtu.be/iTWdjtaSf9g\nAndrew Lukyanenko (Kaggle Grandmaster) \"Getting started with Kaggle\""
  },
  "source": "meta"
}