{
  "id": 341348,
  "title": "Searching an easy way to implement RNN...someone can help me?",
  "url": "/competitions/amex-default-prediction/discussion/341348",
  "author_name": "Zenone",
  "post_date": "2022-08-02T13:37:19.894000",
  "votes": 3,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi.<br>\nCan someone shows me an easy implementation of RNN model (compile, train and predict) given a fixed number of features (n_features)?<br>\nThanks</p>",
  "messages": [
    {
      "id": 1881416,
      "postDate": "2022-08-02T13:37:19.893Z",
      "content": "<p>Hi.<br>\nCan someone shows me an easy implementation of RNN model (compile, train and predict) given a fixed number of features (n_features)?<br>\nThanks</p>",
      "rawMarkdown": "Hi.\nCan someone shows me an easy implementation of RNN model (compile, train and predict) given a fixed number of features (n_features)?\nThanks",
      "votes": 3
    },
    {
      "id": 1883306,
      "postDate": "2022-08-03T18:07:16.973Z",
      "content": "<p>The first and most difficult step is transforming Kaggle's CSV which is 2D (two dimensional) data into 3D data. Discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\" target=\"_blank\">here</a>. Starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a>. Once you have 3D data, then building and training an RNN is only a few lines of code.</p>",
      "rawMarkdown": "The first and most difficult step is transforming Kaggle's CSV which is 2D (two dimensional) data into 3D data. Discussion [here][1]. Starter notebook [here][2]. Once you have 3D data, then building and training an RNN is only a few lines of code.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790",
      "votes": 1,
      "replies": [
        {
          "id": 1884936,
          "postDate": "2022-08-04T18:33:46.773Z",
          "content": "<p>thank you very much<br>\nvery interesting \"pad rows\"</p>",
          "rawMarkdown": "thank you very much\nvery interesting \"pad rows\"",
          "votes": 1
        },
        {
          "id": 1889914,
          "postDate": "2022-08-08T12:41:53.477Z",
          "content": "<p>Just one more question: if I have a train pandas dataframe 2D (596583,178) how can i reshape in 3D?<br>\nI apologize for \"dummy question\" , but it's first time that I work with RNN and i quite disoriented</p>\n<p>Thanks a lot</p>",
          "rawMarkdown": "Just one more question: if I have a train pandas dataframe 2D (596583,178) how can i reshape in 3D?\nI apologize for \"dummy question\" , but it's first time that I work with RNN and i quite disoriented\n\nThanks a lot"
        },
        {
          "id": 1890197,
          "postDate": "2022-08-08T15:33:36.530Z",
          "content": "<p>The first thing you must do is sort the dataframe by customer and date like <code>df = df.sortvalues(['customer','time'])</code>. Next you need to add missing rows to each customer so that each customer has 13 rows. Then finally you convert 2D to 3D with the command <code>data_3d = df.iloc[:, SKIP:].values.reshape((-1, 13, 188))</code> where SKIP skips the columns that are not features like date and customer ID.</p>\n<p>This whole procedure uses a lot of memory, so you will need to convert the dataframe into 3D NumPy arrays in pieces. I prefer to make all the pieces ahead of time and save to disk. But you could make a data loader that converts chunks into 3D during training.</p>",
          "rawMarkdown": "The first thing you must do is sort the dataframe by customer and date like `df = df.sortvalues(['customer','time'])`. Next you need to add missing rows to each customer so that each customer has 13 rows. Then finally you convert 2D to 3D with the command `data_3d = df.iloc[:, SKIP:].values.reshape((-1, 13, 188))` where SKIP skips the columns that are not features like date and customer ID.\n\nThis whole procedure uses a lot of memory, so you will need to convert the dataframe into 3D NumPy arrays in pieces. I prefer to make all the pieces ahead of time and save to disk. But you could make a data loader that converts chunks into 3D during training."
        },
        {
          "id": 1890337,
          "postDate": "2022-08-08T17:15:00.723Z",
          "content": "<p>Thank you very very much </p>",
          "rawMarkdown": "Thank you very very much ",
          "votes": 1
        },
        {
          "id": 1890601,
          "postDate": "2022-08-08T22:11:28.207Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nHello Chris,<br>\nIt's nice to see you participating in the competition. Your answers in the other discussions are absolutely helpful!<br>\nJust have a question, in the XGB Starter notebook you shared, you used KFold to evaluate the model. As that we are investigating a client behavior in the future, why not using TimeSeriesSplit?<br>\nAnd when using your approach for reducing the memory of the Customer_id (using the hex numbers), how could we get the original ones back when preparing our submissions? As that we will add the groupby features to both the training and testing data, we will apply the same operation to both.</p>",
          "rawMarkdown": "@cdeotte \nHello Chris,\nIt's nice to see you participating in the competition. Your answers in the other discussions are absolutely helpful!\nJust have a question, in the XGB Starter notebook you shared, you used KFold to evaluate the model. As that we are investigating a client behavior in the future, why not using TimeSeriesSplit?\nAnd when using your approach for reducing the memory of the Customer_id (using the hex numbers), how could we get the original ones back when preparing our submissions? As that we will add the groupby features to both the training and testing data, we will apply the same operation to both.",
          "votes": 1
        },
        {
          "id": 1890896,
          "postDate": "2022-08-09T05:46:35.280Z",
          "content": "<p>All the targets occur at the same point in time. Therefore there is no need to split the targets with TimeSeriesSplit. </p>\n<p>(In other time series competitions, we may have some January targets, some February targets, some March targets etc. Then we could use TimeSeriesSplit and train with Jan + Feb targets and predict Mar targets. And train with Jan targets and predict Feb + Mar targets etc etc.)</p>\n<p>To get the original hex back. Just make a dataframe that has two columns. One column has customer_id and one column is hex values. Then you can merge this dataframe onto another data frame with only hex to give you customer_id. Or you could merge this dataframe onto customer_id to give you hex.</p>",
          "rawMarkdown": "All the targets occur at the same point in time. Therefore there is no need to split the targets with TimeSeriesSplit. \n\n(In other time series competitions, we may have some January targets, some February targets, some March targets etc. Then we could use TimeSeriesSplit and train with Jan + Feb targets and predict Mar targets. And train with Jan targets and predict Feb + Mar targets etc etc.)\n\nTo get the original hex back. Just make a dataframe that has two columns. One column has customer_id and one column is hex values. Then you can merge this dataframe onto another data frame with only hex to give you customer_id. Or you could merge this dataframe onto customer_id to give you hex.",
          "votes": 2
        },
        {
          "id": 1891043,
          "postDate": "2022-08-09T08:02:22.807Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1905676,
          "postDate": "2022-08-19T08:18:05.130Z",
          "content": "<p>Hi.<br>\nI have some problems installing cupy an d cudf on my desktop (pycharm) even if i have installed CUDA (v11.7) and cuDNN. So I try your notebook direct on Kaggle but when I train the model I have a problem with memory and the notebook stop to run.<br>\nCan you explain me why?</p>\n<p>Thanks a lot</p>",
          "rawMarkdown": "Hi.\nI have some problems installing cupy an d cudf on my desktop (pycharm) even if i have installed CUDA (v11.7) and cuDNN. So I try your notebook direct on Kaggle but when I train the model I have a problem with memory and the notebook stop to run.\nCan you explain me why?\n\nThanks a lot"
        }
      ]
    },
    {
      "id": 1882367,
      "postDate": "2022-08-03T08:35:12.837Z",
      "content": "<p>You may peruse Kaggle notebooks in this regard- </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/raoulma/ny-stock-price-prediction-rnn-lstm-gru\" target=\"_blank\">https://www.kaggle.com/code/raoulma/ny-stock-price-prediction-rnn-lstm-gru</a></li>\n<li><a href=\"https://www.kaggle.com/code/frlemarchand/covid-19-forecasting-with-an-rnn\" target=\"_blank\">https://www.kaggle.com/code/frlemarchand/covid-19-forecasting-with-an-rnn</a></li>\n<li><a href=\"https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series\" target=\"_blank\">https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series</a></li>\n<li><a href=\"https://www.kaggle.com/code/zikazika/using-rnn-and-arima-to-predict-bitcoin-price\" target=\"_blank\">https://www.kaggle.com/code/zikazika/using-rnn-and-arima-to-predict-bitcoin-price</a></li>\n<li><a href=\"https://www.kaggle.com/code/charel/learn-by-example-rnn-lstm-gru-time-series\" target=\"_blank\">https://www.kaggle.com/code/charel/learn-by-example-rnn-lstm-gru-time-series</a></li>\n</ol>\n<p>These are some popular notebooks in this regard. I hope this is useful though!</p>",
      "rawMarkdown": "You may peruse Kaggle notebooks in this regard- \n1. https://www.kaggle.com/code/raoulma/ny-stock-price-prediction-rnn-lstm-gru\n2. https://www.kaggle.com/code/frlemarchand/covid-19-forecasting-with-an-rnn\n3. https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series\n4. https://www.kaggle.com/code/zikazika/using-rnn-and-arima-to-predict-bitcoin-price\n5. https://www.kaggle.com/code/charel/learn-by-example-rnn-lstm-gru-time-series\n\nThese are some popular notebooks in this regard. I hope this is useful though!",
      "votes": 1,
      "replies": [
        {
          "id": 1882793,
          "postDate": "2022-08-03T12:34:05.720Z",
          "content": "<p>Thank you for the links</p>",
          "rawMarkdown": "Thank you for the links"
        }
      ]
    },
    {
      "id": 1881505,
      "postDate": "2022-08-02T14:57:02.700Z",
      "content": "<p>There are a lot of good notebooks if you search for RNN. Maybe these can help:<br>\nTensor flow:<br>\n<a href=\"https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series\" target=\"_blank\">https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series</a><br>\nPytorch:<br>\n<a href=\"https://www.kaggle.com/code/kanncaa1/recurrent-neural-network-with-pytorch\" target=\"_blank\">https://www.kaggle.com/code/kanncaa1/recurrent-neural-network-with-pytorch</a></p>",
      "rawMarkdown": "There are a lot of good notebooks if you search for RNN. Maybe these can help:\nTensor flow:\nhttps://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series\nPytorch:\nhttps://www.kaggle.com/code/kanncaa1/recurrent-neural-network-with-pytorch",
      "votes": 2,
      "replies": [
        {
          "id": 1881721,
          "postDate": "2022-08-02T18:11:57.663Z",
          "content": "<p>Thank you very much</p>",
          "rawMarkdown": "Thank you very much"
        }
      ]
    },
    {
      "id": 1894125,
      "postDate": "2022-08-11T09:34:36.057Z",
      "content": "<p>Hi, this is an example of a notebook I have created to predict a brain stroke<br>\n<a href=\"https://www.kaggle.com/code/jorgeromn/brain-stroke-with-simple-neural-networks-95\" target=\"_blank\">https://www.kaggle.com/code/jorgeromn/brain-stroke-with-simple-neural-networks-95</a></p>",
      "rawMarkdown": "Hi, this is an example of a notebook I have created to predict a brain stroke\n[https://www.kaggle.com/code/jorgeromn/brain-stroke-with-simple-neural-networks-95](https://www.kaggle.com/code/jorgeromn/brain-stroke-with-simple-neural-networks-95)\n"
    }
  ],
  "comments": [
    {
      "id": 1883306,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-08-03T18:07:16.973000",
      "content": "<p>The first and most difficult step is transforming Kaggle's CSV which is 2D (two dimensional) data into 3D data. Discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\" target=\"_blank\">here</a>. Starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a>. Once you have 3D data, then building and training an RNN is only a few lines of code.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1884936,
          "author_name": "Zenone",
          "author_url": "",
          "post_date": "2022-08-04T18:33:46.773000",
          "content": "<p>thank you very much<br>\nvery interesting \"pad rows\"</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1889914,
          "author_name": "Zenone",
          "author_url": "",
          "post_date": "2022-08-08T12:41:53.477000",
          "content": "<p>Just one more question: if I have a train pandas dataframe 2D (596583,178) how can i reshape in 3D?<br>\nI apologize for \"dummy question\" , but it's first time that I work with RNN and i quite disoriented</p>\n<p>Thanks a lot</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1890197,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-08T15:33:36.530000",
          "content": "<p>The first thing you must do is sort the dataframe by customer and date like <code>df = df.sortvalues(['customer','time'])</code>. Next you need to add missing rows to each customer so that each customer has 13 rows. Then finally you convert 2D to 3D with the command <code>data_3d = df.iloc[:, SKIP:].values.reshape((-1, 13, 188))</code> where SKIP skips the columns that are not features like date and customer ID.</p>\n<p>This whole procedure uses a lot of memory, so you will need to convert the dataframe into 3D NumPy arrays in pieces. I prefer to make all the pieces ahead of time and save to disk. But you could make a data loader that converts chunks into 3D during training.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1890337,
          "author_name": "Zenone",
          "author_url": "",
          "post_date": "2022-08-08T17:15:00.723000",
          "content": "<p>Thank you very very much </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1890601,
          "author_name": "Mohamed Eltayeb",
          "author_url": "",
          "post_date": "2022-08-08T22:11:28.207000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <br>\nHello Chris,<br>\nIt's nice to see you participating in the competition. Your answers in the other discussions are absolutely helpful!<br>\nJust have a question, in the XGB Starter notebook you shared, you used KFold to evaluate the model. As that we are investigating a client behavior in the future, why not using TimeSeriesSplit?<br>\nAnd when using your approach for reducing the memory of the Customer_id (using the hex numbers), how could we get the original ones back when preparing our submissions? As that we will add the groupby features to both the training and testing data, we will apply the same operation to both.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1890896,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-09T05:46:35.280000",
          "content": "<p>All the targets occur at the same point in time. Therefore there is no need to split the targets with TimeSeriesSplit. </p>\n<p>(In other time series competitions, we may have some January targets, some February targets, some March targets etc. Then we could use TimeSeriesSplit and train with Jan + Feb targets and predict Mar targets. And train with Jan targets and predict Feb + Mar targets etc etc.)</p>\n<p>To get the original hex back. Just make a dataframe that has two columns. One column has customer_id and one column is hex values. Then you can merge this dataframe onto another data frame with only hex to give you customer_id. Or you could merge this dataframe onto customer_id to give you hex.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1891043,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-09T08:02:22.807000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1905676,
          "author_name": "Zenone",
          "author_url": "",
          "post_date": "2022-08-19T08:18:05.130000",
          "content": "<p>Hi.<br>\nI have some problems installing cupy an d cudf on my desktop (pycharm) even if i have installed CUDA (v11.7) and cuDNN. So I try your notebook direct on Kaggle but when I train the model I have a problem with memory and the notebook stop to run.<br>\nCan you explain me why?</p>\n<p>Thanks a lot</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1882367,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-08-03T08:35:12.837000",
      "content": "<p>You may peruse Kaggle notebooks in this regard- </p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/raoulma/ny-stock-price-prediction-rnn-lstm-gru\" target=\"_blank\">https://www.kaggle.com/code/raoulma/ny-stock-price-prediction-rnn-lstm-gru</a></li>\n<li><a href=\"https://www.kaggle.com/code/frlemarchand/covid-19-forecasting-with-an-rnn\" target=\"_blank\">https://www.kaggle.com/code/frlemarchand/covid-19-forecasting-with-an-rnn</a></li>\n<li><a href=\"https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series\" target=\"_blank\">https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series</a></li>\n<li><a href=\"https://www.kaggle.com/code/zikazika/using-rnn-and-arima-to-predict-bitcoin-price\" target=\"_blank\">https://www.kaggle.com/code/zikazika/using-rnn-and-arima-to-predict-bitcoin-price</a></li>\n<li><a href=\"https://www.kaggle.com/code/charel/learn-by-example-rnn-lstm-gru-time-series\" target=\"_blank\">https://www.kaggle.com/code/charel/learn-by-example-rnn-lstm-gru-time-series</a></li>\n</ol>\n<p>These are some popular notebooks in this regard. I hope this is useful though!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1882793,
          "author_name": "Zenone",
          "author_url": "",
          "post_date": "2022-08-03T12:34:05.720000",
          "content": "<p>Thank you for the links</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1881505,
      "author_name": "Sinclair",
      "author_url": "",
      "post_date": "2022-08-02T14:57:02.700000",
      "content": "<p>There are a lot of good notebooks if you search for RNN. Maybe these can help:<br>\nTensor flow:<br>\n<a href=\"https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series\" target=\"_blank\">https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series</a><br>\nPytorch:<br>\n<a href=\"https://www.kaggle.com/code/kanncaa1/recurrent-neural-network-with-pytorch\" target=\"_blank\">https://www.kaggle.com/code/kanncaa1/recurrent-neural-network-with-pytorch</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1881721,
          "author_name": "Zenone",
          "author_url": "",
          "post_date": "2022-08-02T18:11:57.663000",
          "content": "<p>Thank you very much</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1894125,
      "author_name": "Jorge Román",
      "author_url": "",
      "post_date": "2022-08-11T09:34:36.057000",
      "content": "<p>Hi, this is an example of a notebook I have created to predict a brain stroke<br>\n<a href=\"https://www.kaggle.com/code/jorgeromn/brain-stroke-with-simple-neural-networks-95\" target=\"_blank\">https://www.kaggle.com/code/jorgeromn/brain-stroke-with-simple-neural-networks-95</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1881416": "Hi.\nCan someone shows me an easy implementation of RNN model (compile, train and predict) given a fixed number of features (n_features)?\nThanks",
    "1883306": "The first and most difficult step is transforming Kaggle's CSV which is 2D (two dimensional) data into 3D data. Discussion [here][1]. Starter notebook [here][2]. Once you have 3D data, then building and training an RNN is only a few lines of code.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790",
    "1882367": "You may peruse Kaggle notebooks in this regard- \n1. https://www.kaggle.com/code/raoulma/ny-stock-price-prediction-rnn-lstm-gru\n2. https://www.kaggle.com/code/frlemarchand/covid-19-forecasting-with-an-rnn\n3. https://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series\n4. https://www.kaggle.com/code/zikazika/using-rnn-and-arima-to-predict-bitcoin-price\n5. https://www.kaggle.com/code/charel/learn-by-example-rnn-lstm-gru-time-series\n\nThese are some popular notebooks in this regard. I hope this is useful though!",
    "1881505": "There are a lot of good notebooks if you search for RNN. Maybe these can help:\nTensor flow:\nhttps://www.kaggle.com/code/mayer79/rnn-starter-for-huge-time-series\nPytorch:\nhttps://www.kaggle.com/code/kanncaa1/recurrent-neural-network-with-pytorch",
    "1894125": "Hi, this is an example of a notebook I have created to predict a brain stroke\n[https://www.kaggle.com/code/jorgeromn/brain-stroke-with-simple-neural-networks-95](https://www.kaggle.com/code/jorgeromn/brain-stroke-with-simple-neural-networks-95)\n"
  }
}