{
  "id": 366590,
  "title": "Open Problems | My solution and ideas",
  "url": "/competitions/open-problems-multimodal/discussion/366590",
  "author_name": "Dmitriy Ershov",
  "post_date": "2022-11-16T20:12:26.114000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi! Thanks to all the organizers and kagglers for this interesting, important and hard competition. I want to share my solution and ideas with you because it may be helpful. </p>\n<h1>Multiome</h1>\n<h3>Preprocessing</h3>\n<ol>\n<li>SVD with concatenated <a href=\"https://www.kaggle.com/datasets/stasborodynkin/feature-shop-for-mmscel-multiome\" target=\"_blank\">Feature Shop</a> (best is 256 components)</li>\n<li>SVD with default data (128, 256)</li>\n<li>SVD for targets (best is 128 components)</li>\n<li>Feature selection with <a href=\"https://www.kaggle.com/bejeweled/multiome-rf-feature-selection\" target=\"_blank\">Random Forest</a></li>\n</ol>\n<p>Corr selection, KNN worked bad.</p>\n<h3>Models</h3>\n<ol>\n<li>LGB single regression</li>\n<li>Ridge with pseudo-labelling </li>\n<li>NNs - dense NN with Bi-LSTM, dense NN, 1D CNN, 2D CNN (dense, reshape, 2D convs)</li>\n</ol>\n<p>TabNet, CB worked bad. NNs were trained as single regressions and multi regressions both.</p>\n<h1>CITEseq</h1>\n<h3>Preprocessing</h3>\n<ol>\n<li>Concatenated <a href=\"https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition\" target=\"_blank\">Feature Shop</a> for NN</li>\n<li>SVD with default data (256)</li>\n<li>Feature selection with <a href=\"https://www.kaggle.com/code/bejeweled/siteseq-rf-feature-selection\" target=\"_blank\">Random Forest</a></li>\n<li>Feature selection based on <a href=\"https://www.kaggle.com/code/bejeweled/siteseq-corr-feature-selection\" target=\"_blank\">correlation</a>  -&gt; worked well for single regressions with flipping of negative correlated features</li>\n</ol>\n<h3>Models</h3>\n<ol>\n<li>LGB single regression with RF and corrs features</li>\n<li>CB single regression with RF features</li>\n<li>NNs - dense NN, 1D CNN, 2D CNN, with/without cell embedings</li>\n</ol>\n<p>TabNet also worked bad. </p>\n<h1>Blending</h1>\n<p>I tried 4 different methods:</p>\n<ol>\n<li>Normalized averaging</li>\n<li>Weighted normalized averaging by total oof correlation</li>\n<li>Weighted normalized averaging by single target oof correlation</li>\n<li>Stacking with Ridge</li>\n</ol>\n<p>Third method gives best results.</p>\n<h1>Ideas I did not try or tried a little</h1>\n<ol>\n<li>Transform vectors to distance matrices and fit it to 2D CNN. It may help make connections between features. Also we can build model with dual input - one for feature vector and one for distance matrix</li>\n<li>Augmentations (flip, little noise, vec rotating) for CNN</li>\n<li>Autoencoders</li>\n<li>Clip values and make group of features </li>\n<li>Kalman filter (was too long to compute) </li>\n<li>Fit only non-zero features to NN as sequences</li>\n<li>Different models for different cell types or weighted averaging, because different models show different results on different cell types</li>\n</ol>\n<h1>Some my notebooks with analysis</h1>\n<ol>\n<li>CITEseq targets <a href=\"https://www.kaggle.com/code/bejeweled/mmscel-citeseq-targets-eda?scriptVersionId=109777501\" target=\"_blank\">EDA</a></li>\n<li>Playground <a href=\"https://www.kaggle.com/code/bejeweled/mmscel-all-targs-modeling-playground-fold/notebook\" target=\"_blank\">modeling</a></li>\n<li>CITEseq sklearn different cells <a href=\"https://www.kaggle.com/code/bejeweled/citeseq-sklearn-cells-feature-shop?scriptVersionId=110648831\" target=\"_blank\">modeling</a></li>\n</ol>",
  "messages": [
    {
      "id": 2032792,
      "postDate": "2022-11-16T20:12:26.113Z",
      "content": "<p>Hi! Thanks to all the organizers and kagglers for this interesting, important and hard competition. I want to share my solution and ideas with you because it may be helpful. </p>\n<h1>Multiome</h1>\n<h3>Preprocessing</h3>\n<ol>\n<li>SVD with concatenated <a href=\"https://www.kaggle.com/datasets/stasborodynkin/feature-shop-for-mmscel-multiome\" target=\"_blank\">Feature Shop</a> (best is 256 components)</li>\n<li>SVD with default data (128, 256)</li>\n<li>SVD for targets (best is 128 components)</li>\n<li>Feature selection with <a href=\"https://www.kaggle.com/bejeweled/multiome-rf-feature-selection\" target=\"_blank\">Random Forest</a></li>\n</ol>\n<p>Corr selection, KNN worked bad.</p>\n<h3>Models</h3>\n<ol>\n<li>LGB single regression</li>\n<li>Ridge with pseudo-labelling </li>\n<li>NNs - dense NN with Bi-LSTM, dense NN, 1D CNN, 2D CNN (dense, reshape, 2D convs)</li>\n</ol>\n<p>TabNet, CB worked bad. NNs were trained as single regressions and multi regressions both.</p>\n<h1>CITEseq</h1>\n<h3>Preprocessing</h3>\n<ol>\n<li>Concatenated <a href=\"https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition\" target=\"_blank\">Feature Shop</a> for NN</li>\n<li>SVD with default data (256)</li>\n<li>Feature selection with <a href=\"https://www.kaggle.com/code/bejeweled/siteseq-rf-feature-selection\" target=\"_blank\">Random Forest</a></li>\n<li>Feature selection based on <a href=\"https://www.kaggle.com/code/bejeweled/siteseq-corr-feature-selection\" target=\"_blank\">correlation</a>  -&gt; worked well for single regressions with flipping of negative correlated features</li>\n</ol>\n<h3>Models</h3>\n<ol>\n<li>LGB single regression with RF and corrs features</li>\n<li>CB single regression with RF features</li>\n<li>NNs - dense NN, 1D CNN, 2D CNN, with/without cell embedings</li>\n</ol>\n<p>TabNet also worked bad. </p>\n<h1>Blending</h1>\n<p>I tried 4 different methods:</p>\n<ol>\n<li>Normalized averaging</li>\n<li>Weighted normalized averaging by total oof correlation</li>\n<li>Weighted normalized averaging by single target oof correlation</li>\n<li>Stacking with Ridge</li>\n</ol>\n<p>Third method gives best results.</p>\n<h1>Ideas I did not try or tried a little</h1>\n<ol>\n<li>Transform vectors to distance matrices and fit it to 2D CNN. It may help make connections between features. Also we can build model with dual input - one for feature vector and one for distance matrix</li>\n<li>Augmentations (flip, little noise, vec rotating) for CNN</li>\n<li>Autoencoders</li>\n<li>Clip values and make group of features </li>\n<li>Kalman filter (was too long to compute) </li>\n<li>Fit only non-zero features to NN as sequences</li>\n<li>Different models for different cell types or weighted averaging, because different models show different results on different cell types</li>\n</ol>\n<h1>Some my notebooks with analysis</h1>\n<ol>\n<li>CITEseq targets <a href=\"https://www.kaggle.com/code/bejeweled/mmscel-citeseq-targets-eda?scriptVersionId=109777501\" target=\"_blank\">EDA</a></li>\n<li>Playground <a href=\"https://www.kaggle.com/code/bejeweled/mmscel-all-targs-modeling-playground-fold/notebook\" target=\"_blank\">modeling</a></li>\n<li>CITEseq sklearn different cells <a href=\"https://www.kaggle.com/code/bejeweled/citeseq-sklearn-cells-feature-shop?scriptVersionId=110648831\" target=\"_blank\">modeling</a></li>\n</ol>",
      "rawMarkdown": "Hi! Thanks to all the organizers and kagglers for this interesting, important and hard competition. I want to share my solution and ideas with you because it may be helpful. \n\n# Multiome\n\n### Preprocessing\n\n1. SVD with concatenated [Feature Shop](https://www.kaggle.com/datasets/stasborodynkin/feature-shop-for-mmscel-multiome) (best is 256 components)\n2. SVD with default data (128, 256)\n3. SVD for targets (best is 128 components)\n4. Feature selection with [Random Forest](https://www.kaggle.com/bejeweled/multiome-rf-feature-selection)\n\nCorr selection, KNN worked bad.\n\n### Models\n\n1. LGB single regression\n2. Ridge with pseudo-labelling \n3. NNs - dense NN with Bi-LSTM, dense NN, 1D CNN, 2D CNN (dense, reshape, 2D convs)\n\nTabNet, CB worked bad. NNs were trained as single regressions and multi regressions both.\n\n# CITEseq\n\n### Preprocessing\n\n1. Concatenated [Feature Shop](https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition) for NN\n2. SVD with default data (256)\n3. Feature selection with [Random Forest](https://www.kaggle.com/code/bejeweled/siteseq-rf-feature-selection)\n4. Feature selection based on [correlation](https://www.kaggle.com/code/bejeweled/siteseq-corr-feature-selection)  -> worked well for single regressions with flipping of negative correlated features\n\n### Models\n\n1. LGB single regression with RF and corrs features\n2. CB single regression with RF features\n3. NNs - dense NN, 1D CNN, 2D CNN, with/without cell embedings\n\nTabNet also worked bad. \n\n# Blending\n\nI tried 4 different methods:\n1. Normalized averaging\n2. Weighted normalized averaging by total oof correlation\n3. Weighted normalized averaging by single target oof correlation\n4. Stacking with Ridge\n\nThird method gives best results.\n\n# Ideas I did not try or tried a little\n\n1. Transform vectors to distance matrices and fit it to 2D CNN. It may help make connections between features. Also we can build model with dual input - one for feature vector and one for distance matrix\n2. Augmentations (flip, little noise, vec rotating) for CNN\n3. Autoencoders\n4. Clip values and make group of features \n5. Kalman filter (was too long to compute) \n6. Fit only non-zero features to NN as sequences\n7. Different models for different cell types or weighted averaging, because different models show different results on different cell types\n\n# Some my notebooks with analysis\n\n1. CITEseq targets [EDA](https://www.kaggle.com/code/bejeweled/mmscel-citeseq-targets-eda?scriptVersionId=109777501)\n2. Playground [modeling](https://www.kaggle.com/code/bejeweled/mmscel-all-targs-modeling-playground-fold/notebook)\n3. CITEseq sklearn different cells [modeling](https://www.kaggle.com/code/bejeweled/citeseq-sklearn-cells-feature-shop?scriptVersionId=110648831)",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2032792": "Hi! Thanks to all the organizers and kagglers for this interesting, important and hard competition. I want to share my solution and ideas with you because it may be helpful. \n\n# Multiome\n\n### Preprocessing\n\n1. SVD with concatenated [Feature Shop](https://www.kaggle.com/datasets/stasborodynkin/feature-shop-for-mmscel-multiome) (best is 256 components)\n2. SVD with default data (128, 256)\n3. SVD for targets (best is 128 components)\n4. Feature selection with [Random Forest](https://www.kaggle.com/bejeweled/multiome-rf-feature-selection)\n\nCorr selection, KNN worked bad.\n\n### Models\n\n1. LGB single regression\n2. Ridge with pseudo-labelling \n3. NNs - dense NN with Bi-LSTM, dense NN, 1D CNN, 2D CNN (dense, reshape, 2D convs)\n\nTabNet, CB worked bad. NNs were trained as single regressions and multi regressions both.\n\n# CITEseq\n\n### Preprocessing\n\n1. Concatenated [Feature Shop](https://www.kaggle.com/datasets/alexandervc/feature-shop-for-multimodal-singlecell-competition) for NN\n2. SVD with default data (256)\n3. Feature selection with [Random Forest](https://www.kaggle.com/code/bejeweled/siteseq-rf-feature-selection)\n4. Feature selection based on [correlation](https://www.kaggle.com/code/bejeweled/siteseq-corr-feature-selection)  -> worked well for single regressions with flipping of negative correlated features\n\n### Models\n\n1. LGB single regression with RF and corrs features\n2. CB single regression with RF features\n3. NNs - dense NN, 1D CNN, 2D CNN, with/without cell embedings\n\nTabNet also worked bad. \n\n# Blending\n\nI tried 4 different methods:\n1. Normalized averaging\n2. Weighted normalized averaging by total oof correlation\n3. Weighted normalized averaging by single target oof correlation\n4. Stacking with Ridge\n\nThird method gives best results.\n\n# Ideas I did not try or tried a little\n\n1. Transform vectors to distance matrices and fit it to 2D CNN. It may help make connections between features. Also we can build model with dual input - one for feature vector and one for distance matrix\n2. Augmentations (flip, little noise, vec rotating) for CNN\n3. Autoencoders\n4. Clip values and make group of features \n5. Kalman filter (was too long to compute) \n6. Fit only non-zero features to NN as sequences\n7. Different models for different cell types or weighted averaging, because different models show different results on different cell types\n\n# Some my notebooks with analysis\n\n1. CITEseq targets [EDA](https://www.kaggle.com/code/bejeweled/mmscel-citeseq-targets-eda?scriptVersionId=109777501)\n2. Playground [modeling](https://www.kaggle.com/code/bejeweled/mmscel-all-targs-modeling-playground-fold/notebook)\n3. CITEseq sklearn different cells [modeling](https://www.kaggle.com/code/bejeweled/citeseq-sklearn-cells-feature-shop?scriptVersionId=110648831)"
  }
}