{
  "id": 352208,
  "title": "Anyone using cnn models?",
  "url": "/competitions/open-problems-multimodal/discussion/352208",
  "author_name": "Kuro",
  "post_date": "2022-09-13T12:31:42.977000",
  "votes": 7,
  "comment_count": 17,
  "views": 0,
  "content": "<p>According to this website <a href=\"https://lanceotron.molbiol.ox.ac.uk/projects/peak_search_basic/6243\" target=\"_blank\">https://lanceotron.molbiol.ox.ac.uk/projects/peak_search_basic/6243</a> ，it seems gene GL000XXX.X can be viewed as a 1D curve . Does anyone get good results using cnn ?</p>",
  "messages": [
    {
      "id": 1937340,
      "postDate": "2022-09-13T12:31:42.977Z",
      "content": "<p>According to this website <a href=\"https://lanceotron.molbiol.ox.ac.uk/projects/peak_search_basic/6243\" target=\"_blank\">https://lanceotron.molbiol.ox.ac.uk/projects/peak_search_basic/6243</a> ，it seems gene GL000XXX.X can be viewed as a 1D curve . Does anyone get good results using cnn ?</p>",
      "rawMarkdown": "According to this website https://lanceotron.molbiol.ox.ac.uk/projects/peak_search_basic/6243 ，it seems gene GL000XXX.X can be viewed as a 1D curve . Does anyone get good results using cnn ?",
      "votes": 7
    },
    {
      "id": 2026896,
      "postDate": "2022-11-12T12:09:11.660Z",
      "content": "<p>did you try 1D CNN like this? <a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/202256\" target=\"_blank\">https://www.kaggle.com/competitions/lish-moa/discussion/202256</a>)</p>",
      "rawMarkdown": "did you try 1D CNN like this? https://www.kaggle.com/competitions/lish-moa/discussion/202256)",
      "votes": 1,
      "replies": [
        {
          "id": 2027638,
          "postDate": "2022-11-13T03:30:55.793Z",
          "content": "<p>No ,i didn't. But it sounds interesting.</p>\n<blockquote>\n  <p>CNN structure performs well in feature extraction, but it is rarely used in tabular data because <strong>the correct features ordering is unknown</strong>.<br>\n  A simple idea is to reshape the data directly into a multi-channel image format, and <strong>the correct sorting can be learned</strong> by using FC layer through back propagation.</p>\n</blockquote>\n<p>Thank you for your sharing! I think we need to use a fc layer to rearrange tabular features before feeding into cnn models.</p>",
          "rawMarkdown": "No ,i didn't. But it sounds interesting.\n\n> CNN structure performs well in feature extraction, but it is rarely used in tabular data because **the correct features ordering is unknown**.\nA simple idea is to reshape the data directly into a multi-channel image format, and **the correct sorting can be learned** by using FC layer through back propagation.\n\nThank you for your sharing! I think we need to use a fc layer to rearrange tabular features before feeding into cnn models."
        },
        {
          "id": 2027653,
          "postDate": "2022-11-13T04:11:18.590Z",
          "content": "<p>And i think dimension increasing for input is not working for this competition because our data dim is already so high. Before that we still have to use pca or other method to reduce dim. Then use main components fc projection to form a meaningful image.</p>",
          "rawMarkdown": "And i think dimension increasing for input is not working for this competition because our data dim is already so high. Before that we still have to use pca or other method to reduce dim. Then use main components fc projection to form a meaningful image."
        },
        {
          "id": 2027966,
          "postDate": "2022-11-13T11:47:07.690Z",
          "content": "<p>are u using the important features of which names are correlated with the targets' names? like what are stated in <a href=\"https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras/notebook#Target-normalization\" target=\"_blank\">https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras/notebook#Target-normalization</a>? did you try to group the important features by correlations? somebody said that, keeping those important features out of PCA leads to 0.812…but I haven't try throwing all important features and other features into PCA together…what's your experiences? thanks!</p>",
          "rawMarkdown": "are u using the important features of which names are correlated with the targets' names? like what are stated in https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras/notebook#Target-normalization? did you try to group the important features by correlations? somebody said that, keeping those important features out of PCA leads to 0.812...but I haven't try throwing all important features and other features into PCA together...what's your experiences? thanks!"
        },
        {
          "id": 2028056,
          "postDate": "2022-11-13T13:37:57.697Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2028059,
          "postDate": "2022-11-13T13:42:39.200Z",
          "content": "<p>I've written a simple 1d cnn experiment based on <a href=\"https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras\" target=\"_blank\">All in one : CITEseq &amp; Multiome</a>. It already has important feat &amp; PCA concatenation.</p>\n<p>My groupkfold average cv is about 0.8916 in comparation to AMBROSM's 0.89236. Much better than my previous cnn score.</p>\n<blockquote>\n  <p>my 1d CNN cv:<br>\n  fold0: 0.8885<br>\n  fold1: 0.8952<br>\n  fold2: 0.8913<br>\n  Average  corr = 0.8916</p>\n  <p>AMBROSM's cv :<br>\n   Fold 0:  49 epochs, corr =  0.89033<br>\n  Fold 1:  46 epochs, corr =  0.89536<br>\n  Fold 2:  44 epochs, corr =  0.89140<br>\n  Average  corr = 0.89236</p>\n</blockquote>\n<p>Maybe it's a good model for ensembling.</p>",
          "rawMarkdown": "I've written a simple 1d cnn experiment based on [All in one : CITEseq & Multiome](https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras). It already has important feat & PCA concatenation.\n\nMy groupkfold average cv is about 0.8916 in comparation to AMBROSM's 0.89236. Much better than my previous cnn score.\n\n> my 1d CNN cv:\nfold0: 0.8885\nfold1: 0.8952\nfold2: 0.8913\nAverage  corr = 0.8916\n\n>AMBROSM's cv :\n Fold 0:  49 epochs, corr =  0.89033\nFold 1:  46 epochs, corr =  0.89536\nFold 2:  44 epochs, corr =  0.89140\nAverage  corr = 0.89236\n\nMaybe it's a good model for ensembling."
        },
        {
          "id": 2028226,
          "postDate": "2022-11-13T16:28:30.210Z",
          "content": "<p>Thanks! But it seems that the average correlation of 1d CNN cv as 0.8916 is smaller than the mean correlation cv of \"all in one\" experiments which is 0.8939…though a higher cv on training dataset does not mean a higher cv on testing dataset…so may I know your cv of the public submission? Thanks! By the way, why are you comparing your cv to AMBROSM's…is he at the first place right now?</p>",
          "rawMarkdown": "Thanks! But it seems that the average correlation of 1d CNN cv as 0.8916 is smaller than the mean correlation cv of \"all in one\" experiments which is 0.8939...though a higher cv on training dataset does not mean a higher cv on testing dataset...so may I know your cv of the public submission? Thanks! By the way, why are you comparing your cv to AMBROSM's...is he at the first place right now?"
        },
        {
          "id": 2028246,
          "postDate": "2022-11-13T16:47:34.403Z",
          "content": "<p>I just test if 1d cnn works and I don't think I've further time for hyper param tunning😦. You know there are much more structures for cnn than mlp. <br>\nActually, I've ensembled citeseq part above to my submission and boost my lb a little.</p>",
          "rawMarkdown": "I just test if 1d cnn works and I don't think I've further time for hyper param tunning😦. You know there are much more structures for cnn than mlp. \nActually, I've ensembled citeseq part above to my submission and boost my lb a little."
        },
        {
          "id": 2028288,
          "postDate": "2022-11-13T17:54:28.500Z",
          "content": "<p>thanks for your reply!</p>",
          "rawMarkdown": "thanks for your reply!"
        },
        {
          "id": 2028365,
          "postDate": "2022-11-13T20:50:04.440Z",
          "content": "<p>but may I know which datasets you are using to ensemble your citeseq part? from your own models or public datasets? Thanks!</p>",
          "rawMarkdown": "but may I know which datasets you are using to ensemble your citeseq part? from your own models or public datasets? Thanks!"
        },
        {
          "id": 2028369,
          "postDate": "2022-11-13T20:52:44.777Z",
          "content": "<p>looks great, hoping to see your sharing post after the competition end</p>",
          "rawMarkdown": "looks great, hoping to see your sharing post after the competition end"
        },
        {
          "id": 2028537,
          "postDate": "2022-11-14T03:24:40.890Z",
          "content": "<p><a href=\"https://www.kaggle.com/xushankaggle\" target=\"_blank\">@xushankaggle</a> to my best ensembled submission.</p>",
          "rawMarkdown": "@xushankaggle to my best ensembled submission."
        }
      ]
    },
    {
      "id": 1937346,
      "postDate": "2022-09-13T12:38:20.567Z",
      "content": "<p>I have tried , but didn't get better results than a simple mlp baseline .</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7955478%2F951463dc083d1b7f8c73938323e16889%2FWB%20Chart%202022_9_13%2020_35_55%20(1).png?generation=1663072686079619&amp;alt=media\" alt=\"\"></p>\n<p>And this is my v12 source code modified from <a href=\"https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors#Simple-Model:-MLP\" target=\"_blank\">this book</a>. I just change the mlp class.</p>\n<pre><code>class SublayerConnection(nn.Module):\n    \"\"\"A residual connection followed by a layer norm.\n    Note for code simplicity the norm is first as opposed to last.\n    \"\"\"\n    def __init__(self, sublayer ,skip_connection_weight=1.0):\n        super(SublayerConnection, self).__init__()\n        # self.dropout = nn.Dropout(dropout)\n        self.sublayer = sublayer\n        self.skip_connection_weight = skip_connection_weight\n    def forward(self, x):\n        \"Apply residual connection to any sublayer with the same size.\"\n        sublayer = self.sublayer\n        return x + (sublayer(x) * self.skip_connection_weight)\n\nclass MLP(nn.Module):\n    def __init__(self, layer_size_lst, add_final_activation=False):\n        super().__init__()\n        self.max_pooling = nn.MaxPool1d(2, stride=2)\n        self.fc = nn.Linear(48*3 , layer_size_lst[-1])\n        layer_lst = [48]*10\n        layers = [\n            nn.Conv1d(1, layer_lst[0], 3, stride=1, padding=1,padding_mode='zeros'),\n            nn.ReLU(),\n            nn.MaxPool1d(3, stride=3),\n            ]\n        for i in range(1,len(layer_lst)):\n            layers += nn.Sequential(\n                SublayerConnection(nn.Sequential(\n                    nn.Conv1d(layer_lst[i-1], layer_lst[i], 5, stride=1, padding=2,padding_mode='zeros'),\n                    nn.BatchNorm1d(layer_lst[i]),\n                    nn.ReLU(),\n                ),1.0),\n                nn.MaxPool1d(3, stride=3),\n            )\n        self.convolutions = nn.Sequential(*layers)\n    def forward(self, x):\n        if not g.d1:\n            x = x.to_dense()\n        x = x.view(x.shape[0], 1,-1) \n        x = x \n        x = self.convolutions(x)\n        x = x.view(x.shape[0], -1)\n        x = self.fc(x)\n        return x\n</code></pre>",
      "rawMarkdown": "I have tried , but didn't get better results than a simple mlp baseline .\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7955478%2F951463dc083d1b7f8c73938323e16889%2FWB%20Chart%202022_9_13%2020_35_55%20(1).png?generation=1663072686079619&alt=media)\n\n\nAnd this is my v12 source code modified from [this book](https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors#Simple-Model:-MLP). I just change the mlp class.\n\n```python\nclass SublayerConnection(nn.Module):\n    \"\"\"A residual connection followed by a layer norm.\n    Note for code simplicity the norm is first as opposed to last.\n    \"\"\"\n    def __init__(self, sublayer ,skip_connection_weight=1.0):\n        super(SublayerConnection, self).__init__()\n        # self.dropout = nn.Dropout(dropout)\n        self.sublayer = sublayer\n        self.skip_connection_weight = skip_connection_weight\n    def forward(self, x):\n        \"Apply residual connection to any sublayer with the same size.\"\n        sublayer = self.sublayer\n        return x + (sublayer(x) * self.skip_connection_weight)\n\nclass MLP(nn.Module):\n    def __init__(self, layer_size_lst, add_final_activation=False):\n        super().__init__()\n        self.max_pooling = nn.MaxPool1d(2, stride=2)\n        self.fc = nn.Linear(48*3 , layer_size_lst[-1])\n        layer_lst = [48]*10\n        layers = [\n            nn.Conv1d(1, layer_lst[0], 3, stride=1, padding=1,padding_mode='zeros'),\n            nn.ReLU(),\n            nn.MaxPool1d(3, stride=3),\n            ]\n        for i in range(1,len(layer_lst)):\n            layers += nn.Sequential(\n                SublayerConnection(nn.Sequential(\n                    nn.Conv1d(layer_lst[i-1], layer_lst[i], 5, stride=1, padding=2,padding_mode='zeros'),\n                    nn.BatchNorm1d(layer_lst[i]),\n                    nn.ReLU(),\n                ),1.0),\n                nn.MaxPool1d(3, stride=3),\n            )\n        self.convolutions = nn.Sequential(*layers)\n    def forward(self, x):\n        if not g.d1:\n            x = x.to_dense()\n        x = x.view(x.shape[0], 1,-1) \n        x = x \n        x = self.convolutions(x)\n        x = x.view(x.shape[0], -1)\n        x = self.fc(x)\n        return x\n\n```",
      "votes": 2
    },
    {
      "id": 2028371,
      "postDate": "2022-11-13T20:55:02.130Z",
      "content": "<p>I tried end-2-end big cnn models to raw data in cite month ago, but it didn't work well, though I thought it is interesting idea at that time. I think the cnn after feature engineering might be more useful and good to ensemble (though there is too limited time for me to try it)</p>",
      "rawMarkdown": "I tried end-2-end big cnn models to raw data in cite month ago, but it didn't work well, though I thought it is interesting idea at that time. I think the cnn after feature engineering might be more useful and good to ensemble (though there is too limited time for me to try it)"
    },
    {
      "id": 1937558,
      "postDate": "2022-09-13T14:53:13.193Z",
      "content": "<p>I tried using a Longformer with a window size of 500 but it did not work better than a MLP baseline either.</p>\n<p>I suppose CNN also does not work because there exists correlation between inputs that are far apart..</p>",
      "rawMarkdown": "I tried using a Longformer with a window size of 500 but it did not work better than a MLP baseline either.\n\nI suppose CNN also does not work because there exists correlation between inputs that are far apart..",
      "replies": [
        {
          "id": 1958833,
          "postDate": "2022-09-27T16:49:03.610Z",
          "content": "<p>Hi, I'm also thinking about selfattention recently.Can you share some details about your transformer model?</p>",
          "rawMarkdown": "Hi, I'm also thinking about selfattention recently.Can you share some details about your transformer model?"
        }
      ]
    },
    {
      "id": 1939086,
      "postDate": "2022-09-14T14:18:54.583Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2026896,
      "author_name": "XuShanKaggle",
      "author_url": "",
      "post_date": "2022-11-12T12:09:11.660000",
      "content": "<p>did you try 1D CNN like this? <a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/202256\" target=\"_blank\">https://www.kaggle.com/competitions/lish-moa/discussion/202256</a>)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2027638,
          "author_name": "Kuro",
          "author_url": "",
          "post_date": "2022-11-13T03:30:55.793000",
          "content": "<p>No ,i didn't. But it sounds interesting.</p>\n<blockquote>\n  <p>CNN structure performs well in feature extraction, but it is rarely used in tabular data because <strong>the correct features ordering is unknown</strong>.<br>\n  A simple idea is to reshape the data directly into a multi-channel image format, and <strong>the correct sorting can be learned</strong> by using FC layer through back propagation.</p>\n</blockquote>\n<p>Thank you for your sharing! I think we need to use a fc layer to rearrange tabular features before feeding into cnn models.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2027653,
          "author_name": "Kuro",
          "author_url": "",
          "post_date": "2022-11-13T04:11:18.590000",
          "content": "<p>And i think dimension increasing for input is not working for this competition because our data dim is already so high. Before that we still have to use pca or other method to reduce dim. Then use main components fc projection to form a meaningful image.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2027966,
          "author_name": "XuShanKaggle",
          "author_url": "",
          "post_date": "2022-11-13T11:47:07.690000",
          "content": "<p>are u using the important features of which names are correlated with the targets' names? like what are stated in <a href=\"https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras/notebook#Target-normalization\" target=\"_blank\">https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras/notebook#Target-normalization</a>? did you try to group the important features by correlations? somebody said that, keeping those important features out of PCA leads to 0.812…but I haven't try throwing all important features and other features into PCA together…what's your experiences? thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2028056,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-11-13T13:37:57.697000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2028059,
          "author_name": "Kuro",
          "author_url": "",
          "post_date": "2022-11-13T13:42:39.200000",
          "content": "<p>I've written a simple 1d cnn experiment based on <a href=\"https://www.kaggle.com/code/pourchot/all-in-one-citeseq-multiome-with-keras\" target=\"_blank\">All in one : CITEseq &amp; Multiome</a>. It already has important feat &amp; PCA concatenation.</p>\n<p>My groupkfold average cv is about 0.8916 in comparation to AMBROSM's 0.89236. Much better than my previous cnn score.</p>\n<blockquote>\n  <p>my 1d CNN cv:<br>\n  fold0: 0.8885<br>\n  fold1: 0.8952<br>\n  fold2: 0.8913<br>\n  Average  corr = 0.8916</p>\n  <p>AMBROSM's cv :<br>\n   Fold 0:  49 epochs, corr =  0.89033<br>\n  Fold 1:  46 epochs, corr =  0.89536<br>\n  Fold 2:  44 epochs, corr =  0.89140<br>\n  Average  corr = 0.89236</p>\n</blockquote>\n<p>Maybe it's a good model for ensembling.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2028226,
          "author_name": "XuShanKaggle",
          "author_url": "",
          "post_date": "2022-11-13T16:28:30.210000",
          "content": "<p>Thanks! But it seems that the average correlation of 1d CNN cv as 0.8916 is smaller than the mean correlation cv of \"all in one\" experiments which is 0.8939…though a higher cv on training dataset does not mean a higher cv on testing dataset…so may I know your cv of the public submission? Thanks! By the way, why are you comparing your cv to AMBROSM's…is he at the first place right now?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2028246,
          "author_name": "Kuro",
          "author_url": "",
          "post_date": "2022-11-13T16:47:34.403000",
          "content": "<p>I just test if 1d cnn works and I don't think I've further time for hyper param tunning😦. You know there are much more structures for cnn than mlp. <br>\nActually, I've ensembled citeseq part above to my submission and boost my lb a little.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2028288,
          "author_name": "XuShanKaggle",
          "author_url": "",
          "post_date": "2022-11-13T17:54:28.500000",
          "content": "<p>thanks for your reply!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2028365,
          "author_name": "XuShanKaggle",
          "author_url": "",
          "post_date": "2022-11-13T20:50:04.440000",
          "content": "<p>but may I know which datasets you are using to ensemble your citeseq part? from your own models or public datasets? Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2028369,
          "author_name": "no-magic",
          "author_url": "",
          "post_date": "2022-11-13T20:52:44.777000",
          "content": "<p>looks great, hoping to see your sharing post after the competition end</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2028537,
          "author_name": "Kuro",
          "author_url": "",
          "post_date": "2022-11-14T03:24:40.890000",
          "content": "<p><a href=\"https://www.kaggle.com/xushankaggle\" target=\"_blank\">@xushankaggle</a> to my best ensembled submission.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1937346,
      "author_name": "Kuro",
      "author_url": "",
      "post_date": "2022-09-13T12:38:20.567000",
      "content": "<p>I have tried , but didn't get better results than a simple mlp baseline .</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7955478%2F951463dc083d1b7f8c73938323e16889%2FWB%20Chart%202022_9_13%2020_35_55%20(1).png?generation=1663072686079619&amp;alt=media\" alt=\"\"></p>\n<p>And this is my v12 source code modified from <a href=\"https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors#Simple-Model:-MLP\" target=\"_blank\">this book</a>. I just change the mlp class.</p>\n<pre><code>class SublayerConnection(nn.Module):\n    \"\"\"A residual connection followed by a layer norm.\n    Note for code simplicity the norm is first as opposed to last.\n    \"\"\"\n    def __init__(self, sublayer ,skip_connection_weight=1.0):\n        super(SublayerConnection, self).__init__()\n        # self.dropout = nn.Dropout(dropout)\n        self.sublayer = sublayer\n        self.skip_connection_weight = skip_connection_weight\n    def forward(self, x):\n        \"Apply residual connection to any sublayer with the same size.\"\n        sublayer = self.sublayer\n        return x + (sublayer(x) * self.skip_connection_weight)\n\nclass MLP(nn.Module):\n    def __init__(self, layer_size_lst, add_final_activation=False):\n        super().__init__()\n        self.max_pooling = nn.MaxPool1d(2, stride=2)\n        self.fc = nn.Linear(48*3 , layer_size_lst[-1])\n        layer_lst = [48]*10\n        layers = [\n            nn.Conv1d(1, layer_lst[0], 3, stride=1, padding=1,padding_mode='zeros'),\n            nn.ReLU(),\n            nn.MaxPool1d(3, stride=3),\n            ]\n        for i in range(1,len(layer_lst)):\n            layers += nn.Sequential(\n                SublayerConnection(nn.Sequential(\n                    nn.Conv1d(layer_lst[i-1], layer_lst[i], 5, stride=1, padding=2,padding_mode='zeros'),\n                    nn.BatchNorm1d(layer_lst[i]),\n                    nn.ReLU(),\n                ),1.0),\n                nn.MaxPool1d(3, stride=3),\n            )\n        self.convolutions = nn.Sequential(*layers)\n    def forward(self, x):\n        if not g.d1:\n            x = x.to_dense()\n        x = x.view(x.shape[0], 1,-1) \n        x = x \n        x = self.convolutions(x)\n        x = x.view(x.shape[0], -1)\n        x = self.fc(x)\n        return x\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2028371,
      "author_name": "no-magic",
      "author_url": "",
      "post_date": "2022-11-13T20:55:02.130000",
      "content": "<p>I tried end-2-end big cnn models to raw data in cite month ago, but it didn't work well, though I thought it is interesting idea at that time. I think the cnn after feature engineering might be more useful and good to ensemble (though there is too limited time for me to try it)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1937558,
      "author_name": "pinouche",
      "author_url": "",
      "post_date": "2022-09-13T14:53:13.193000",
      "content": "<p>I tried using a Longformer with a window size of 500 but it did not work better than a MLP baseline either.</p>\n<p>I suppose CNN also does not work because there exists correlation between inputs that are far apart..</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1958833,
          "author_name": "Kuro",
          "author_url": "",
          "post_date": "2022-09-27T16:49:03.610000",
          "content": "<p>Hi, I'm also thinking about selfattention recently.Can you share some details about your transformer model?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1939086,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-14T14:18:54.583000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1937340": "According to this website https://lanceotron.molbiol.ox.ac.uk/projects/peak_search_basic/6243 ，it seems gene GL000XXX.X can be viewed as a 1D curve . Does anyone get good results using cnn ?",
    "2026896": "did you try 1D CNN like this? https://www.kaggle.com/competitions/lish-moa/discussion/202256)",
    "1937346": "I have tried , but didn't get better results than a simple mlp baseline .\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7955478%2F951463dc083d1b7f8c73938323e16889%2FWB%20Chart%202022_9_13%2020_35_55%20(1).png?generation=1663072686079619&alt=media)\n\n\nAnd this is my v12 source code modified from [this book](https://www.kaggle.com/code/fabiencrom/msci-multiome-torch-quickstart-w-sparse-tensors#Simple-Model:-MLP). I just change the mlp class.\n\n```python\nclass SublayerConnection(nn.Module):\n    \"\"\"A residual connection followed by a layer norm.\n    Note for code simplicity the norm is first as opposed to last.\n    \"\"\"\n    def __init__(self, sublayer ,skip_connection_weight=1.0):\n        super(SublayerConnection, self).__init__()\n        # self.dropout = nn.Dropout(dropout)\n        self.sublayer = sublayer\n        self.skip_connection_weight = skip_connection_weight\n    def forward(self, x):\n        \"Apply residual connection to any sublayer with the same size.\"\n        sublayer = self.sublayer\n        return x + (sublayer(x) * self.skip_connection_weight)\n\nclass MLP(nn.Module):\n    def __init__(self, layer_size_lst, add_final_activation=False):\n        super().__init__()\n        self.max_pooling = nn.MaxPool1d(2, stride=2)\n        self.fc = nn.Linear(48*3 , layer_size_lst[-1])\n        layer_lst = [48]*10\n        layers = [\n            nn.Conv1d(1, layer_lst[0], 3, stride=1, padding=1,padding_mode='zeros'),\n            nn.ReLU(),\n            nn.MaxPool1d(3, stride=3),\n            ]\n        for i in range(1,len(layer_lst)):\n            layers += nn.Sequential(\n                SublayerConnection(nn.Sequential(\n                    nn.Conv1d(layer_lst[i-1], layer_lst[i], 5, stride=1, padding=2,padding_mode='zeros'),\n                    nn.BatchNorm1d(layer_lst[i]),\n                    nn.ReLU(),\n                ),1.0),\n                nn.MaxPool1d(3, stride=3),\n            )\n        self.convolutions = nn.Sequential(*layers)\n    def forward(self, x):\n        if not g.d1:\n            x = x.to_dense()\n        x = x.view(x.shape[0], 1,-1) \n        x = x \n        x = self.convolutions(x)\n        x = x.view(x.shape[0], -1)\n        x = self.fc(x)\n        return x\n\n```",
    "2028371": "I tried end-2-end big cnn models to raw data in cite month ago, but it didn't work well, though I thought it is interesting idea at that time. I think the cnn after feature engineering might be more useful and good to ensemble (though there is too limited time for me to try it)",
    "1937558": "I tried using a Longformer with a window size of 500 but it did not work better than a MLP baseline either.\n\nI suppose CNN also does not work because there exists correlation between inputs that are far apart..",
    "1939086": ""
  }
}