{
  "id": 240277,
  "title": "On the time-step dimension ",
  "url": "/competitions/seti-breakthrough-listen/discussion/240277",
  "author_name": "Baran Hashemi",
  "post_date": "2021-05-19T06:17:23.627000",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am a little bit confused by the kernels. And I think it is important for some type of models. As far as I understood the time-step dimension is 273 in (6, 273, 256).  First of all, am I right?<br>\nSecondly, If I am right, then why some kernels use the 256 dimensions for the stacking, or if they want to use channel-wise models they transpose the model to (6, 256, 273) (e.g. the biGRU kernel <a href=\"url\" target=\"_blank\">https://www.kaggle.com/xhlulu/seti-bidirectional-gru-in-keras</a>)?<br>\nIn total, how important do you think it is to mind which sequence of dimensions one uses?</p>",
  "messages": [
    {
      "id": 1314373,
      "postDate": "2021-05-19T06:17:23.627Z",
      "content": "<p>I am a little bit confused by the kernels. And I think it is important for some type of models. As far as I understood the time-step dimension is 273 in (6, 273, 256).  First of all, am I right?<br>\nSecondly, If I am right, then why some kernels use the 256 dimensions for the stacking, or if they want to use channel-wise models they transpose the model to (6, 256, 273) (e.g. the biGRU kernel <a href=\"url\" target=\"_blank\">https://www.kaggle.com/xhlulu/seti-bidirectional-gru-in-keras</a>)?<br>\nIn total, how important do you think it is to mind which sequence of dimensions one uses?</p>",
      "rawMarkdown": "I am a little bit confused by the kernels. And I think it is important for some type of models. As far as I understood the time-step dimension is 273 in (6, 273, 256).  First of all, am I right?\nSecondly, If I am right, then why some kernels use the 256 dimensions for the stacking, or if they want to use channel-wise models they transpose the model to (6, 256, 273) (e.g. the biGRU kernel [https://www.kaggle.com/xhlulu/seti-bidirectional-gru-in-keras](url))?\nIn total, how important do you think it is to mind which sequence of dimensions one uses?",
      "votes": 6
    },
    {
      "id": 1318346,
      "postDate": "2021-05-22T08:32:43.237Z",
      "content": "<p>Guys, I encountered sth very strange here. And apparently, I am not the only one. <br>\nWhen I use any RNN unit on top of either the image itself or its embedding, and I use 273 or 1638 as the time-step dimension (seq dim as in Pytorch), I get a very bad result (CV 50+). However, when I use 256 as the time-step dimension, the results are much better (CV 90+). I really don't understand why.</p>",
      "rawMarkdown": "Guys, I encountered sth very strange here. And apparently, I am not the only one. \nWhen I use any RNN unit on top of either the image itself or its embedding, and I use 273 or 1638 as the time-step dimension (seq dim as in Pytorch), I get a very bad result (CV 50+). However, when I use 256 as the time-step dimension, the results are much better (CV 90+). I really don't understand why.\n",
      "votes": 3
    },
    {
      "id": 1318769,
      "postDate": "2021-05-22T15:03:13.817Z",
      "content": "<p>It looks quite logical, that you have better results when using 256 as time-step dimension. Take two slices of image containing target: one 1638x16 and other 16x256. It's much easies to check the slice contains target if you see full time picture even in narrow bandwidth range.</p>",
      "rawMarkdown": "It looks quite logical, that you have better results when using 256 as time-step dimension. Take two slices of image containing target: one 1638x16 and other 16x256. It's much easies to check the slice contains target if you see full time picture even in narrow bandwidth range.\n",
      "votes": 1,
      "replies": [
        {
          "id": 1319016,
          "postDate": "2021-05-22T19:21:56.747Z",
          "content": "<p>Hmm, sorry but i don't still get it. You say \"<em>It's much easies to check the slice contains target if you see full time picture even in narrow bandwidth range</em>.\" But when i state 1638 as the time dim., the RNN unit is looking at the full time picture. It is like the RNN is looking at each slice of time with freq. Vector as its feature vector. Just to make sure we are in the same page, 1638 ( or 273) is the time dim and 256 is the freq. dim.</p>",
          "rawMarkdown": "Hmm, sorry but i don't still get it. You say \"*It's much easies to check the slice contains target if you see full time picture even in narrow bandwidth range*.\" But when i state 1638 as the time dim., the RNN unit is looking at the full time picture. It is like the RNN is looking at each slice of time with freq. Vector as its feature vector. Just to make sure we are in the same page, 1638 ( or 273) is the time dim and 256 is the freq. dim."
        },
        {
          "id": 1319237,
          "postDate": "2021-05-23T03:45:23.840Z",
          "content": "<p>Yes, we are on the same page (1638 = 273*6 - it is time dimension). RNN have full access to information during training, it can actually (in theory) use memory to store what is important. But… training an RNN (to use information that used to be 273 time steps before) is not an easy task. On the other hand (using 256 as a time step) RNN has almost all the necessary information on one page and does not need to strain.</p>",
          "rawMarkdown": "Yes, we are on the same page (1638 = 273*6 - it is time dimension). RNN have full access to information during training, it can actually (in theory) use memory to store what is important. But... training an RNN (to use information that used to be 273 time steps before) is not an easy task. On the other hand (using 256 as a time step) RNN has almost all the necessary information on one page and does not need to strain."
        },
        {
          "id": 1319423,
          "postDate": "2021-05-23T07:14:10.340Z",
          "content": "<p>Tnx for clarification. Let me see if I understood correctly. You are basically saying that 1638 is too long for RNN, and since if one uses the feature vectors as the time steps, theoretically there should be no problem, hence RNN works better.<br>\nBut, even when one uses 273 as the time step and run an RNN for each channel, one gets a poor result (and if one does it other way around 256 as the time step and 273 as feature vectors the performance gets high again)<br>\nBtw, do you have any references for your argument regarding RNN's compatibility when using transposed information?</p>",
          "rawMarkdown": "Tnx for clarification. Let me see if I understood correctly. You are basically saying that 1638 is too long for RNN, and since if one uses the feature vectors as the time steps, theoretically there should be no problem, hence RNN works better.\nBut, even when one uses 273 as the time step and run an RNN for each channel, one gets a poor result (and if one does it other way around 256 as the time step and 273 as feature vectors the performance gets high again)\nBtw, do you have any references for your argument regarding RNN's compatibility when using transposed information?"
        },
        {
          "id": 1319621,
          "postDate": "2021-05-23T11:33:24.087Z",
          "content": "<p>My two cents on this: When you use the time dimension (273) as the time dimension in the RNN (sorry for the wording 😊) the model \"sees\" (=gets as input) all the frequencies (256) at one point in time, on the other hand if you use the frequencies as the time dimension, the model's input is one frequence over the whole time span (273).<br>\nIt may just be that the latter incorporates more useful information for the model to learn from (carrying this information through the network).</p>",
          "rawMarkdown": "My two cents on this: When you use the time dimension (273) as the time dimension in the RNN (sorry for the wording 😊) the model \"sees\" (=gets as input) all the frequencies (256) at one point in time, on the other hand if you use the frequencies as the time dimension, the model's input is one frequence over the whole time span (273).\nIt may just be that the latter incorporates more useful information for the model to learn from (carrying this information through the network).",
          "votes": 1
        },
        {
          "id": 1320131,
          "postDate": "2021-05-23T18:59:14.513Z",
          "content": "<p>Hmmm, yeah it seems you are right. Oddly, I have seen no paper using freq. as the sequence dim. in any similar model in any context.<br>\n<a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> <a href=\"https://www.kaggle.com/achikin\" target=\"_blank\">@achikin</a>  Do you know of any references or examples that people used word embeddings instead of word sequences (or any other example in any other context) as the actual sequence?<br>\n<a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> what is your opinion on this?</p>",
          "rawMarkdown": "Hmmm, yeah it seems you are right. Oddly, I have seen no paper using freq. as the sequence dim. in any similar model in any context.\n@hannes82 @achikin  Do you know of any references or examples that people used word embeddings instead of word sequences (or any other example in any other context) as the actual sequence?\n@xhlulu what is your opinion on this?"
        },
        {
          "id": 1321019,
          "postDate": "2021-05-24T13:33:15.457Z",
          "content": "<p><a href=\"https://www.kaggle.com/rythian47\" target=\"_blank\">@rythian47</a> No I don't know any. Not sure if this would work for NLP as it is everything about context there.</p>",
          "rawMarkdown": "@rythian47 No I don't know any. Not sure if this would work for NLP as it is everything about context there."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1318346,
      "author_name": "Baran Hashemi",
      "author_url": "",
      "post_date": "2021-05-22T08:32:43.237000",
      "content": "<p>Guys, I encountered sth very strange here. And apparently, I am not the only one. <br>\nWhen I use any RNN unit on top of either the image itself or its embedding, and I use 273 or 1638 as the time-step dimension (seq dim as in Pytorch), I get a very bad result (CV 50+). However, when I use 256 as the time-step dimension, the results are much better (CV 90+). I really don't understand why.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1318769,
      "author_name": "Anton Chikin",
      "author_url": "",
      "post_date": "2021-05-22T15:03:13.817000",
      "content": "<p>It looks quite logical, that you have better results when using 256 as time-step dimension. Take two slices of image containing target: one 1638x16 and other 16x256. It's much easies to check the slice contains target if you see full time picture even in narrow bandwidth range.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1319016,
          "author_name": "Baran Hashemi",
          "author_url": "",
          "post_date": "2021-05-22T19:21:56.747000",
          "content": "<p>Hmm, sorry but i don't still get it. You say \"<em>It's much easies to check the slice contains target if you see full time picture even in narrow bandwidth range</em>.\" But when i state 1638 as the time dim., the RNN unit is looking at the full time picture. It is like the RNN is looking at each slice of time with freq. Vector as its feature vector. Just to make sure we are in the same page, 1638 ( or 273) is the time dim and 256 is the freq. dim.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1319237,
          "author_name": "Anton Chikin",
          "author_url": "",
          "post_date": "2021-05-23T03:45:23.840000",
          "content": "<p>Yes, we are on the same page (1638 = 273*6 - it is time dimension). RNN have full access to information during training, it can actually (in theory) use memory to store what is important. But… training an RNN (to use information that used to be 273 time steps before) is not an easy task. On the other hand (using 256 as a time step) RNN has almost all the necessary information on one page and does not need to strain.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1319423,
          "author_name": "Baran Hashemi",
          "author_url": "",
          "post_date": "2021-05-23T07:14:10.340000",
          "content": "<p>Tnx for clarification. Let me see if I understood correctly. You are basically saying that 1638 is too long for RNN, and since if one uses the feature vectors as the time steps, theoretically there should be no problem, hence RNN works better.<br>\nBut, even when one uses 273 as the time step and run an RNN for each channel, one gets a poor result (and if one does it other way around 256 as the time step and 273 as feature vectors the performance gets high again)<br>\nBtw, do you have any references for your argument regarding RNN's compatibility when using transposed information?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1319621,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-05-23T11:33:24.087000",
          "content": "<p>My two cents on this: When you use the time dimension (273) as the time dimension in the RNN (sorry for the wording 😊) the model \"sees\" (=gets as input) all the frequencies (256) at one point in time, on the other hand if you use the frequencies as the time dimension, the model's input is one frequence over the whole time span (273).<br>\nIt may just be that the latter incorporates more useful information for the model to learn from (carrying this information through the network).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1320131,
          "author_name": "Baran Hashemi",
          "author_url": "",
          "post_date": "2021-05-23T18:59:14.513000",
          "content": "<p>Hmmm, yeah it seems you are right. Oddly, I have seen no paper using freq. as the sequence dim. in any similar model in any context.<br>\n<a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> <a href=\"https://www.kaggle.com/achikin\" target=\"_blank\">@achikin</a>  Do you know of any references or examples that people used word embeddings instead of word sequences (or any other example in any other context) as the actual sequence?<br>\n<a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> what is your opinion on this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1321019,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-05-24T13:33:15.457000",
          "content": "<p><a href=\"https://www.kaggle.com/rythian47\" target=\"_blank\">@rythian47</a> No I don't know any. Not sure if this would work for NLP as it is everything about context there.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1314373": "I am a little bit confused by the kernels. And I think it is important for some type of models. As far as I understood the time-step dimension is 273 in (6, 273, 256).  First of all, am I right?\nSecondly, If I am right, then why some kernels use the 256 dimensions for the stacking, or if they want to use channel-wise models they transpose the model to (6, 256, 273) (e.g. the biGRU kernel [https://www.kaggle.com/xhlulu/seti-bidirectional-gru-in-keras](url))?\nIn total, how important do you think it is to mind which sequence of dimensions one uses?",
    "1318346": "Guys, I encountered sth very strange here. And apparently, I am not the only one. \nWhen I use any RNN unit on top of either the image itself or its embedding, and I use 273 or 1638 as the time-step dimension (seq dim as in Pytorch), I get a very bad result (CV 50+). However, when I use 256 as the time-step dimension, the results are much better (CV 90+). I really don't understand why.\n",
    "1318769": "It looks quite logical, that you have better results when using 256 as time-step dimension. Take two slices of image containing target: one 1638x16 and other 16x256. It's much easies to check the slice contains target if you see full time picture even in narrow bandwidth range.\n"
  }
}