{
  "id": 366420,
  "title": "my summary(72nd)",
  "url": "/competitions/open-problems-multimodal/writeups/ym-my-summary-72nd",
  "author_name": "",
  "post_date": "2022-11-30T12:12:34.813Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thank you all very much.<br>\nI share my ingenuity in this task.</p>\n<p>Citeseq：<br>\nI did not PCA the explanatory variables and used the explanatory variables with higher std than GAPDH (housekeeping gene) for prediction.<br>\nThe dropout rate was set to 0.75 in the first layer of the MLP to automate the selection of explanatory variables.<br>\nCatboost made predictions for each of the 140 targets; Catboost was more accurate than lightgbm without hyperparameter tuning.</p>\n<p>Multi：<br>\nexplanatory variables were dimensionally compressed using SVD for each chromosome. The number of explanatory variables for each chromosome was compressed to the number of explanatory variables divided by 20.<br>\nObjective variables were converted to int8. This is expected not only to compress the data but also to round measurement error.</p>\n<p>Thanks to Dr.Alexander Chervov！</p>\n<p>お疲れ様でした。</p>\n<p>私が今回のタスクで行った取り組みを共有致します。</p>\n<p>Citeseq:<br>\n・説明変数はＰＣＡ実施せず、GAPDH(House keeping gene)よりstdが大きい説明変数を予測に使用した。<br>\n・MLPの第一層目でDropout rateを0.75とかなり大きな値を取り説明変数の選択を自動化した。<br>\n・Catboostは140ターゲット一つずつ予測を行った。ハイパーパラメータチューニングしない条件ではlightgbmと比較しcatboostは予測精度が高い。</p>\n<p>Multi:<br>\n・説明変数はchromosome毎にSVDを用いて次元圧縮した。各Chromosomeの説明変数の数を20で<br>\n割った数に圧縮した。<br>\n・目的変数をint8に変換した。これはデータの圧縮のみでなく測定誤差を丸め込めることも期待した。</p>\n<p>Alexander Chervovさんありがとうございました！</p>",
  "messages": [
    {
      "id": "2031524",
      "postDate": "11/16/2022 06:17:09",
      "content": "<p>Thank you all very much.<br>\nI share my ingenuity in this task.</p>\n<p>Citeseq：<br>\nI did not PCA the explanatory variables and used the explanatory variables with higher std than GAPDH (housekeeping gene) for prediction.<br>\nThe dropout rate was set to 0.75 in the first layer of the MLP to automate the selection of explanatory variables.<br>\nCatboost made predictions for each of the 140 targets; Catboost was more accurate than lightgbm without hyperparameter tuning.</p>\n<p>Multi：<br>\nexplanatory variables were dimensionally compressed using SVD for each chromosome. The number of explanatory variables for each chromosome was compressed to the number of explanatory variables divided by 20.<br>\nObjective variables were converted to int8. This is expected not only to compress the data but also to round measurement error.</p>\n<p>Thanks to Dr.Alexander Chervov！</p>\n<p>お疲れ様でした。</p>\n<p>私が今回のタスクで行った取り組みを共有致します。</p>\n<p>Citeseq:<br>\n・説明変数はＰＣＡ実施せず、GAPDH(House keeping gene)よりstdが大きい説明変数を予測に使用した。<br>\n・MLPの第一層目でDropout rateを0.75とかなり大きな値を取り説明変数の選択を自動化した。<br>\n・Catboostは140ターゲット一つずつ予測を行った。ハイパーパラメータチューニングしない条件ではlightgbmと比較しcatboostは予測精度が高い。</p>\n<p>Multi:<br>\n・説明変数はchromosome毎にSVDを用いて次元圧縮した。各Chromosomeの説明変数の数を20で<br>\n割った数に圧縮した。<br>\n・目的変数をint8に変換した。これはデータの圧縮のみでなく測定誤差を丸め込めることも期待した。</p>\n<p>Alexander Chervovさんありがとうございました！</p>",
      "rawMarkdown": "Thank you all very much.\nI share my ingenuity in this task.\n\nCiteseq：\nI did not PCA the explanatory variables and used the explanatory variables with higher std than GAPDH (housekeeping gene) for prediction.\nThe dropout rate was set to 0.75 in the first layer of the MLP to automate the selection of explanatory variables.\nCatboost made predictions for each of the 140 targets; Catboost was more accurate than lightgbm without hyperparameter tuning.\n\nMulti：\nexplanatory variables were dimensionally compressed using SVD for each chromosome. The number of explanatory variables for each chromosome was compressed to the number of explanatory variables divided by 20.\nObjective variables were converted to int8. This is expected not only to compress the data but also to round measurement error.\n\nThanks to Dr.Alexander Chervov！\n\nお疲れ様でした。\n\n私が今回のタスクで行った取り組みを共有致します。\n\nCiteseq:\n・説明変数はＰＣＡ実施せず、GAPDH(House keeping gene)よりstdが大きい説明変数を予測に使用した。\n・MLPの第一層目でDropout rateを0.75とかなり大きな値を取り説明変数の選択を自動化した。\n・Catboostは140ターゲット一つずつ予測を行った。ハイパーパラメータチューニングしない条件ではlightgbmと比較しcatboostは予測精度が高い。\n\nMulti:\n・説明変数はchromosome毎にSVDを用いて次元圧縮した。各Chromosomeの説明変数の数を20で\n割った数に圧縮した。\n・目的変数をint8に変換した。これはデータの圧縮のみでなく測定誤差を丸め込めることも期待した。\n\nAlexander Chervovさんありがとうございました！",
      "votes": null
    },
    {
      "id": "2044914",
      "postDate": "11/26/2022 20:00:38",
      "content": "<p>Congratulations with the medal and insightful approach ! <br>\n(And thanks for mentioning me).</p>\n<p>May I ask you about the following:</p>\n<p>It would be great - if you can share importances for features - if you have results in that direction.<br>\nWe plan to make further analysis of the data from the biological perspective - if you would have time to take part - that would be honor for us.<br>\nPS<br>\nTechnically: may be create some like \"Kaggle dataset\" and save csv files with importances <br>\nSomething like that: <a href=\"https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances\" target=\"_blank\">https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances</a></p>",
      "rawMarkdown": "Congratulations with the medal and insightful approach ! \n(And thanks for mentioning me).\n\nMay I ask you about the following:\n\nIt would be great - if you can share importances for features - if you have results in that direction.\nWe plan to make further analysis of the data from the biological perspective - if you would have time to take part - that would be honor for us.\nPS\nTechnically: may be create some like \"Kaggle dataset\" and save csv files with importances \nSomething like that: https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances",
      "votes": null
    },
    {
      "id": "2049427",
      "postDate": "11/30/2022 06:25:05",
      "content": "<p>Thank you very much!!!<br>\nI will share the importance of catboost predicted for single target of cite task.<br>\nKeras' permutation importance is not enough for my calculator resources. I gave up on it…</p>\n<p><a href=\"https://www.kaggle.com/datasets/miyawakiyoshifumi/cite-cat-importance\" target=\"_blank\">https://www.kaggle.com/datasets/miyawakiyoshifumi/cite-cat-importance</a></p>",
      "rawMarkdown": "Thank you very much!!!\nI will share the importance of catboost predicted for single target of cite task.\nKeras' permutation importance is not enough for my calculator resources. I gave up on it...\n\nhttps://www.kaggle.com/datasets/miyawakiyoshifumi/cite-cat-importance",
      "votes": null
    },
    {
      "id": "2051526",
      "postDate": "12/01/2022 13:36:41",
      "content": "<p>Thank you very much !</p>\n<p>May I kindly ask to add a few words of descriptions of files,<br>\nand is there any possibility to provide .csv or .txt file - just feature name and importance ? </p>\n<p>PS<br>\nWe are now trying to relate biology to features importances.<br>\nSome notebooks: <br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/mmscel-protein-analysis\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/mmscel-protein-analysis</a><br>\n<a href=\"https://www.kaggle.com/code/dimagromyko/gse148127-lasso-engineering\" target=\"_blank\">https://www.kaggle.com/code/dimagromyko/gse148127-lasso-engineering</a><br>\netc…</p>",
      "rawMarkdown": "Thank you very much !\n\nMay I kindly ask to add a few words of descriptions of files,\nand is there any possibility to provide .csv or .txt file - just feature name and importance ? \n\nPS\nWe are now trying to relate biology to features importances.\nSome notebooks: \nhttps://www.kaggle.com/code/antoninadolgorukova/mmscel-protein-analysis\nhttps://www.kaggle.com/code/dimagromyko/gse148127-lasso-engineering\netc...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2044914,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "11/26/2022 20:00:38",
      "content": "<p>Congratulations with the medal and insightful approach ! <br>\n(And thanks for mentioning me).</p>\n<p>May I ask you about the following:</p>\n<p>It would be great - if you can share importances for features - if you have results in that direction.<br>\nWe plan to make further analysis of the data from the biological perspective - if you would have time to take part - that would be honor for us.<br>\nPS<br>\nTechnically: may be create some like \"Kaggle dataset\" and save csv files with importances <br>\nSomething like that: <a href=\"https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances\" target=\"_blank\">https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2049427,
          "author_name": "miyawakiyoshifumi",
          "author_url": "",
          "post_date": "11/30/2022 06:25:05",
          "content": "<p>Thank you very much!!!<br>\nI will share the importance of catboost predicted for single target of cite task.<br>\nKeras' permutation importance is not enough for my calculator resources. I gave up on it…</p>\n<p><a href=\"https://www.kaggle.com/datasets/miyawakiyoshifumi/cite-cat-importance\" target=\"_blank\">https://www.kaggle.com/datasets/miyawakiyoshifumi/cite-cat-importance</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2051526,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "12/01/2022 13:36:41",
          "content": "<p>Thank you very much !</p>\n<p>May I kindly ask to add a few words of descriptions of files,<br>\nand is there any possibility to provide .csv or .txt file - just feature name and importance ? </p>\n<p>PS<br>\nWe are now trying to relate biology to features importances.<br>\nSome notebooks: <br>\n<a href=\"https://www.kaggle.com/code/antoninadolgorukova/mmscel-protein-analysis\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/mmscel-protein-analysis</a><br>\n<a href=\"https://www.kaggle.com/code/dimagromyko/gse148127-lasso-engineering\" target=\"_blank\">https://www.kaggle.com/code/dimagromyko/gse148127-lasso-engineering</a><br>\netc…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2031524": "Thank you all very much.\nI share my ingenuity in this task.\n\nCiteseq：\nI did not PCA the explanatory variables and used the explanatory variables with higher std than GAPDH (housekeeping gene) for prediction.\nThe dropout rate was set to 0.75 in the first layer of the MLP to automate the selection of explanatory variables.\nCatboost made predictions for each of the 140 targets; Catboost was more accurate than lightgbm without hyperparameter tuning.\n\nMulti：\nexplanatory variables were dimensionally compressed using SVD for each chromosome. The number of explanatory variables for each chromosome was compressed to the number of explanatory variables divided by 20.\nObjective variables were converted to int8. This is expected not only to compress the data but also to round measurement error.\n\nThanks to Dr.Alexander Chervov！\n\nお疲れ様でした。\n\n私が今回のタスクで行った取り組みを共有致します。\n\nCiteseq:\n・説明変数はＰＣＡ実施せず、GAPDH(House keeping gene)よりstdが大きい説明変数を予測に使用した。\n・MLPの第一層目でDropout rateを0.75とかなり大きな値を取り説明変数の選択を自動化した。\n・Catboostは140ターゲット一つずつ予測を行った。ハイパーパラメータチューニングしない条件ではlightgbmと比較しcatboostは予測精度が高い。\n\nMulti:\n・説明変数はchromosome毎にSVDを用いて次元圧縮した。各Chromosomeの説明変数の数を20で\n割った数に圧縮した。\n・目的変数をint8に変換した。これはデータの圧縮のみでなく測定誤差を丸め込めることも期待した。\n\nAlexander Chervovさんありがとうございました！",
    "2044914": "Congratulations with the medal and insightful approach ! \n(And thanks for mentioning me).\n\nMay I ask you about the following:\n\nIt would be great - if you can share importances for features - if you have results in that direction.\nWe plan to make further analysis of the data from the biological perspective - if you would have time to take part - that would be honor for us.\nPS\nTechnically: may be create some like \"Kaggle dataset\" and save csv files with importances \nSomething like that: https://www.kaggle.com/datasets/kaggledummie007/msci-cite-importances",
    "2049427": "Thank you very much!!!\nI will share the importance of catboost predicted for single target of cite task.\nKeras' permutation importance is not enough for my calculator resources. I gave up on it...\n\nhttps://www.kaggle.com/datasets/miyawakiyoshifumi/cite-cat-importance",
    "2051526": "Thank you very much !\n\nMay I kindly ask to add a few words of descriptions of files,\nand is there any possibility to provide .csv or .txt file - just feature name and importance ? \n\nPS\nWe are now trying to relate biology to features importances.\nSome notebooks: \nhttps://www.kaggle.com/code/antoninadolgorukova/mmscel-protein-analysis\nhttps://www.kaggle.com/code/dimagromyko/gse148127-lasso-engineering\netc..."
  },
  "source": "meta"
}