{
  "id": 521778,
  "title": "How to get Train and Test split data ",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/521778",
  "author_name": "",
  "post_date": "2024-07-22T20:11:24.235975900Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>After going over the solution from <a href=\"https://www.kaggle.com/code/abhinavsuri/anatomy-image-visualization-overview-rsna-raids\" target=\"_blank\">https://www.kaggle.com/code/abhinavsuri/anatomy-image-visualization-overview-rsna-raids</a> . It gives  two variables data  \"coor_entries\" and \"df_coor\". </p>\n<p>coor_entries = df_coor[df_coor['study_id'] == int(patient['study_id'])]</p>\n<p>and</p>\n<p>df_coor = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv')</p>\n<p>I am trying to split the data using train_test_split() function . I have passed both  \"coor_entries\" and \"df_coor\" into the function and trying to split the data and return back to X_train, X_test, y_train, y_test. When i do that  i got \"ValueError\"</p>\n<p>X_train, X_test, y_train, y_test = train_test_split(np.array(coor_entries), df_coor , test_size=0.1, random_state=42,stratify=df_coor)   # Complete the code to split the data with test_size as 0.1</p>\n<p>Error:<br>\nValueError: Found input variables with inconsistent numbers of samples: [25, 48692]</p>\n<p>Question: How can i split the train data into  X_train, X_test, y_train, y_test ? So, once i got X_train, X_test, y_train, y_test, I can do the model building as defined below </p>\n<p>I appreciate your help . Thank you</p>\n<h1>Encoding the target labels:</h1>\n<p>enc = LabelBinarizer()                                        <br>\ny_train_encoded = enc.fit_transform(y_train)        <br>\ny_test_encoded=enc.transform(y_test)                  </p>\n<h1>Data Normalization:</h1>\n<p>X_train_normalized = X_train.astype('float32')/255.0<br>\nX_test_normalized = X_test.astype('float32')/255.0</p>\n<h1>Model Building:</h1>\n<h1>Clearing backend</h1>\n<p>backend.clear_session()</p>\n<h1>Fixing the seed for random number generators</h1>\n<p>np.random.seed(42)<br>\nrandom.seed(42)<br>\ntf.random.set_seed(42)</p>\n<h1>Intializing a sequential model</h1>\n<p>model1 = Sequential()                             </p>\n<h1>Complete the code to add the first conv layer with 128 filters and kernel size 3x3 , padding 'same' provides the output size same as the input size</h1>\n<h1>Input_shape denotes input image dimension of images</h1>\n<p>model1.add(Conv2D(128, (3, 3), activation='relu', padding=\"same\", input_shape=(64, 64, 3)))</p>\n<h1>Complete the code to add the max pooling to reduce the size of output of first conv layer</h1>\n<p>model1.add(MaxPooling2D((2, 2), padding = 'same'))</p>\n<h1>Complete the code to create two similar convolution and max-pooling layers activation = relu</h1>\n<p>model1.add(Conv2D(64, (3, 3), activation='relu', padding=\"same\"))<br>\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))</p>\n<p>model1.add(Conv2D(32, (3, 3), activation='relu', padding=\"same\"))<br>\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))</p>\n<h1>Complete the code to flatten the output of the conv layer after max pooling to make it ready for creating dense connections</h1>\n<p>model1.add(Flatten())</p>\n<h1>Complete the code to add a fully connected dense layer with 16 neurons</h1>\n<p>model1.add(Dense(16, activation='relu'))<br>\nmodel1.add(Dropout(0.3))</p>\n<h1>Complete the code to add the output layer with 12 neurons and activation functions as softmax since this is a multi-class classification problem</h1>\n<p>model1.add(Dense(12, activation='softmax'))</p>\n<h1>Complete the code to use the Adam Optimizer</h1>\n<p>opt=Adam()</p>\n<h1>Complete the code to Compile the model using suitable metric for loss fucntion</h1>\n<p>model1.compile(optimizer=opt, loss='categorical_crossentropy', metrics=['accuracy'])</p>\n<h1>Complete the code to generate the summary of the model</h1>\n<p>model1.summary()</p>\n<p>history_1 = model1.fit(<br>\n            X_train_normalized, y_train_encoded,<br>\n            epochs=30,<br>\n            validation_data=(X_test_normalized,y_test_encoded),<br>\n            batch_size=32,<br>\n            verbose=2<br>\n)</p>\n<p>plt.plot(history_1.history['accuracy'])<br>\nplt.plot(history_1.history['test_accuracy'])<br>\nplt.title('Model Accuracy')<br>\nplt.ylabel('Accuracy')<br>\nplt.xlabel('Epoch')<br>\nplt.legend(['Train', 'Test'], loc='upper left')<br>\nplt.show()</p>",
  "messages": [
    {
      "id": "2932338",
      "postDate": "07/22/2024 20:11:24",
      "content": "<p>After going over the solution from <a href=\"https://www.kaggle.com/code/abhinavsuri/anatomy-image-visualization-overview-rsna-raids\" target=\"_blank\">https://www.kaggle.com/code/abhinavsuri/anatomy-image-visualization-overview-rsna-raids</a> . It gives  two variables data  \"coor_entries\" and \"df_coor\". </p>\n<p>coor_entries = df_coor[df_coor['study_id'] == int(patient['study_id'])]</p>\n<p>and</p>\n<p>df_coor = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv')</p>\n<p>I am trying to split the data using train_test_split() function . I have passed both  \"coor_entries\" and \"df_coor\" into the function and trying to split the data and return back to X_train, X_test, y_train, y_test. When i do that  i got \"ValueError\"</p>\n<p>X_train, X_test, y_train, y_test = train_test_split(np.array(coor_entries), df_coor , test_size=0.1, random_state=42,stratify=df_coor)   # Complete the code to split the data with test_size as 0.1</p>\n<p>Error:<br>\nValueError: Found input variables with inconsistent numbers of samples: [25, 48692]</p>\n<p>Question: How can i split the train data into  X_train, X_test, y_train, y_test ? So, once i got X_train, X_test, y_train, y_test, I can do the model building as defined below </p>\n<p>I appreciate your help . Thank you</p>\n<h1>Encoding the target labels:</h1>\n<p>enc = LabelBinarizer()                                        <br>\ny_train_encoded = enc.fit_transform(y_train)        <br>\ny_test_encoded=enc.transform(y_test)                  </p>\n<h1>Data Normalization:</h1>\n<p>X_train_normalized = X_train.astype('float32')/255.0<br>\nX_test_normalized = X_test.astype('float32')/255.0</p>\n<h1>Model Building:</h1>\n<h1>Clearing backend</h1>\n<p>backend.clear_session()</p>\n<h1>Fixing the seed for random number generators</h1>\n<p>np.random.seed(42)<br>\nrandom.seed(42)<br>\ntf.random.set_seed(42)</p>\n<h1>Intializing a sequential model</h1>\n<p>model1 = Sequential()                             </p>\n<h1>Complete the code to add the first conv layer with 128 filters and kernel size 3x3 , padding 'same' provides the output size same as the input size</h1>\n<h1>Input_shape denotes input image dimension of images</h1>\n<p>model1.add(Conv2D(128, (3, 3), activation='relu', padding=\"same\", input_shape=(64, 64, 3)))</p>\n<h1>Complete the code to add the max pooling to reduce the size of output of first conv layer</h1>\n<p>model1.add(MaxPooling2D((2, 2), padding = 'same'))</p>\n<h1>Complete the code to create two similar convolution and max-pooling layers activation = relu</h1>\n<p>model1.add(Conv2D(64, (3, 3), activation='relu', padding=\"same\"))<br>\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))</p>\n<p>model1.add(Conv2D(32, (3, 3), activation='relu', padding=\"same\"))<br>\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))</p>\n<h1>Complete the code to flatten the output of the conv layer after max pooling to make it ready for creating dense connections</h1>\n<p>model1.add(Flatten())</p>\n<h1>Complete the code to add a fully connected dense layer with 16 neurons</h1>\n<p>model1.add(Dense(16, activation='relu'))<br>\nmodel1.add(Dropout(0.3))</p>\n<h1>Complete the code to add the output layer with 12 neurons and activation functions as softmax since this is a multi-class classification problem</h1>\n<p>model1.add(Dense(12, activation='softmax'))</p>\n<h1>Complete the code to use the Adam Optimizer</h1>\n<p>opt=Adam()</p>\n<h1>Complete the code to Compile the model using suitable metric for loss fucntion</h1>\n<p>model1.compile(optimizer=opt, loss='categorical_crossentropy', metrics=['accuracy'])</p>\n<h1>Complete the code to generate the summary of the model</h1>\n<p>model1.summary()</p>\n<p>history_1 = model1.fit(<br>\n            X_train_normalized, y_train_encoded,<br>\n            epochs=30,<br>\n            validation_data=(X_test_normalized,y_test_encoded),<br>\n            batch_size=32,<br>\n            verbose=2<br>\n)</p>\n<p>plt.plot(history_1.history['accuracy'])<br>\nplt.plot(history_1.history['test_accuracy'])<br>\nplt.title('Model Accuracy')<br>\nplt.ylabel('Accuracy')<br>\nplt.xlabel('Epoch')<br>\nplt.legend(['Train', 'Test'], loc='upper left')<br>\nplt.show()</p>",
      "rawMarkdown": "After going over the solution from https://www.kaggle.com/code/abhinavsuri/anatomy-image-visualization-overview-rsna-raids . It gives  two variables data  \"coor_entries\" and \"df_coor\". \n\ncoor_entries = df_coor[df_coor['study_id'] == int(patient['study_id'])]\n\nand\n\ndf_coor = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv')\n\nI am trying to split the data using train_test_split() function . I have passed both  \"coor_entries\" and \"df_coor\" into the function and trying to split the data and return back to X_train, X_test, y_train, y_test. When i do that  i got \"ValueError\"\n\nX_train, X_test, y_train, y_test = train_test_split(np.array(coor_entries), df_coor , test_size=0.1, random_state=42,stratify=df_coor)   # Complete the code to split the data with test_size as 0.1\n\nError:\nValueError: Found input variables with inconsistent numbers of samples: [25, 48692]\n\nQuestion: How can i split the train data into  X_train, X_test, y_train, y_test ? So, once i got X_train, X_test, y_train, y_test, I can do the model building as defined below \n\nI appreciate your help . Thank you\n\n#Encoding the target labels:\nenc = LabelBinarizer()                                        \ny_train_encoded = enc.fit_transform(y_train)        \ny_test_encoded=enc.transform(y_test)                  \n\n#Data Normalization:\n X_train_normalized = X_train.astype('float32')/255.0\nX_test_normalized = X_test.astype('float32')/255.0\n\n#Model Building:\n# Clearing backend\nbackend.clear_session()\n# Fixing the seed for random number generators\nnp.random.seed(42)\nrandom.seed(42)\ntf.random.set_seed(42)\n\n# Intializing a sequential model\nmodel1 = Sequential()                             \n\n# Complete the code to add the first conv layer with 128 filters and kernel size 3x3 , padding 'same' provides the output size same as the input size\n# Input_shape denotes input image dimension of images\nmodel1.add(Conv2D(128, (3, 3), activation='relu', padding=\"same\", input_shape=(64, 64, 3)))\n\n# Complete the code to add the max pooling to reduce the size of output of first conv layer\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))\n\n# Complete the code to create two similar convolution and max-pooling layers activation = relu\nmodel1.add(Conv2D(64, (3, 3), activation='relu', padding=\"same\"))\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))\n\nmodel1.add(Conv2D(32, (3, 3), activation='relu', padding=\"same\"))\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))\n\n# Complete the code to flatten the output of the conv layer after max pooling to make it ready for creating dense connections\nmodel1.add(Flatten())\n\n# Complete the code to add a fully connected dense layer with 16 neurons\nmodel1.add(Dense(16, activation='relu'))\nmodel1.add(Dropout(0.3))\n# Complete the code to add the output layer with 12 neurons and activation functions as softmax since this is a multi-class classification problem\nmodel1.add(Dense(12, activation='softmax'))\n\n# Complete the code to use the Adam Optimizer\nopt=Adam()\n# Complete the code to Compile the model using suitable metric for loss fucntion\nmodel1.compile(optimizer=opt, loss='categorical_crossentropy', metrics=['accuracy'])\n\n# Complete the code to generate the summary of the model\nmodel1.summary()\n\nhistory_1 = model1.fit(\n            X_train_normalized, y_train_encoded,\n            epochs=30,\n            validation_data=(X_test_normalized,y_test_encoded),\n            batch_size=32,\n            verbose=2\n)\n\nplt.plot(history_1.history['accuracy'])\nplt.plot(history_1.history['test_accuracy'])\nplt.title('Model Accuracy')\nplt.ylabel('Accuracy')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Test'], loc='upper left')\nplt.show()",
      "votes": null
    },
    {
      "id": "2932534",
      "postDate": "07/23/2024 05:05:09",
      "content": "<p>From my understanding, <code>patient</code> is a single patient's record: <code>patient = train.iloc[1]</code>.<br>\nSo that's why you get <code>ValueError: Found input variables with inconsistent numbers of samples: [25, 48692]</code> because <code>coor_entries</code> is 25 records long and <code>df_coor</code> is 48692.</p>\n<p><code>train_test_split(np.array(coor_entries), df_coor)</code> does not make any sense. What are the features and labels you are trying to get?</p>\n<p>I am guessing you want to split <code>df_coor</code> by <code>study_id</code> groups. Here is one option:</p>\n<pre><code>all_study_ids = df_coor.study_id.unique()\ntrain_study_ids, test_study_ids = train_test_split(all_study_ids, test_size=, random_state=)\ntrain_df_coor = df_coor[df_coor.study_id.isin(train_study_ids)]\ntest_df_coor = df_coor[df_coor.study_id.isin(test_study_ids)]\n</code></pre>\n<p>These would be your labels. The paths to each study would be your features, given that you are taking some image as input.</p>",
      "rawMarkdown": "From my understanding, `patient` is a single patient's record: `patient = train.iloc[1]`.\nSo that's why you get `ValueError: Found input variables with inconsistent numbers of samples: [25, 48692]` because `coor_entries` is 25 records long and `df_coor` is 48692.\n\n`train_test_split(np.array(coor_entries), df_coor)` does not make any sense. What are the features and labels you are trying to get?\n\nI am guessing you want to split `df_coor` by `study_id` groups. Here is one option:\n```python\nall_study_ids = df_coor.study_id.unique()\ntrain_study_ids, test_study_ids = train_test_split(all_study_ids, test_size=0.1, random_state=42)\ntrain_df_coor = df_coor[df_coor.study_id.isin(train_study_ids)]\ntest_df_coor = df_coor[df_coor.study_id.isin(test_study_ids)]\n```\nThese would be your labels. The paths to each study would be your features, given that you are taking some image as input.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2932534,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "07/23/2024 05:05:09",
      "content": "<p>From my understanding, <code>patient</code> is a single patient's record: <code>patient = train.iloc[1]</code>.<br>\nSo that's why you get <code>ValueError: Found input variables with inconsistent numbers of samples: [25, 48692]</code> because <code>coor_entries</code> is 25 records long and <code>df_coor</code> is 48692.</p>\n<p><code>train_test_split(np.array(coor_entries), df_coor)</code> does not make any sense. What are the features and labels you are trying to get?</p>\n<p>I am guessing you want to split <code>df_coor</code> by <code>study_id</code> groups. Here is one option:</p>\n<pre><code>all_study_ids = df_coor.study_id.unique()\ntrain_study_ids, test_study_ids = train_test_split(all_study_ids, test_size=, random_state=)\ntrain_df_coor = df_coor[df_coor.study_id.isin(train_study_ids)]\ntest_df_coor = df_coor[df_coor.study_id.isin(test_study_ids)]\n</code></pre>\n<p>These would be your labels. The paths to each study would be your features, given that you are taking some image as input.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2932338": "After going over the solution from https://www.kaggle.com/code/abhinavsuri/anatomy-image-visualization-overview-rsna-raids . It gives  two variables data  \"coor_entries\" and \"df_coor\". \n\ncoor_entries = df_coor[df_coor['study_id'] == int(patient['study_id'])]\n\nand\n\ndf_coor = pd.read_csv('/kaggle/input/rsna-2024-lumbar-spine-degenerative-classification/train_label_coordinates.csv')\n\nI am trying to split the data using train_test_split() function . I have passed both  \"coor_entries\" and \"df_coor\" into the function and trying to split the data and return back to X_train, X_test, y_train, y_test. When i do that  i got \"ValueError\"\n\nX_train, X_test, y_train, y_test = train_test_split(np.array(coor_entries), df_coor , test_size=0.1, random_state=42,stratify=df_coor)   # Complete the code to split the data with test_size as 0.1\n\nError:\nValueError: Found input variables with inconsistent numbers of samples: [25, 48692]\n\nQuestion: How can i split the train data into  X_train, X_test, y_train, y_test ? So, once i got X_train, X_test, y_train, y_test, I can do the model building as defined below \n\nI appreciate your help . Thank you\n\n#Encoding the target labels:\nenc = LabelBinarizer()                                        \ny_train_encoded = enc.fit_transform(y_train)        \ny_test_encoded=enc.transform(y_test)                  \n\n#Data Normalization:\n X_train_normalized = X_train.astype('float32')/255.0\nX_test_normalized = X_test.astype('float32')/255.0\n\n#Model Building:\n# Clearing backend\nbackend.clear_session()\n# Fixing the seed for random number generators\nnp.random.seed(42)\nrandom.seed(42)\ntf.random.set_seed(42)\n\n# Intializing a sequential model\nmodel1 = Sequential()                             \n\n# Complete the code to add the first conv layer with 128 filters and kernel size 3x3 , padding 'same' provides the output size same as the input size\n# Input_shape denotes input image dimension of images\nmodel1.add(Conv2D(128, (3, 3), activation='relu', padding=\"same\", input_shape=(64, 64, 3)))\n\n# Complete the code to add the max pooling to reduce the size of output of first conv layer\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))\n\n# Complete the code to create two similar convolution and max-pooling layers activation = relu\nmodel1.add(Conv2D(64, (3, 3), activation='relu', padding=\"same\"))\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))\n\nmodel1.add(Conv2D(32, (3, 3), activation='relu', padding=\"same\"))\nmodel1.add(MaxPooling2D((2, 2), padding = 'same'))\n\n# Complete the code to flatten the output of the conv layer after max pooling to make it ready for creating dense connections\nmodel1.add(Flatten())\n\n# Complete the code to add a fully connected dense layer with 16 neurons\nmodel1.add(Dense(16, activation='relu'))\nmodel1.add(Dropout(0.3))\n# Complete the code to add the output layer with 12 neurons and activation functions as softmax since this is a multi-class classification problem\nmodel1.add(Dense(12, activation='softmax'))\n\n# Complete the code to use the Adam Optimizer\nopt=Adam()\n# Complete the code to Compile the model using suitable metric for loss fucntion\nmodel1.compile(optimizer=opt, loss='categorical_crossentropy', metrics=['accuracy'])\n\n# Complete the code to generate the summary of the model\nmodel1.summary()\n\nhistory_1 = model1.fit(\n            X_train_normalized, y_train_encoded,\n            epochs=30,\n            validation_data=(X_test_normalized,y_test_encoded),\n            batch_size=32,\n            verbose=2\n)\n\nplt.plot(history_1.history['accuracy'])\nplt.plot(history_1.history['test_accuracy'])\nplt.title('Model Accuracy')\nplt.ylabel('Accuracy')\nplt.xlabel('Epoch')\nplt.legend(['Train', 'Test'], loc='upper left')\nplt.show()",
    "2932534": "From my understanding, `patient` is a single patient's record: `patient = train.iloc[1]`.\nSo that's why you get `ValueError: Found input variables with inconsistent numbers of samples: [25, 48692]` because `coor_entries` is 25 records long and `df_coor` is 48692.\n\n`train_test_split(np.array(coor_entries), df_coor)` does not make any sense. What are the features and labels you are trying to get?\n\nI am guessing you want to split `df_coor` by `study_id` groups. Here is one option:\n```python\nall_study_ids = df_coor.study_id.unique()\ntrain_study_ids, test_study_ids = train_test_split(all_study_ids, test_size=0.1, random_state=42)\ntrain_df_coor = df_coor[df_coor.study_id.isin(train_study_ids)]\ntest_df_coor = df_coor[df_coor.study_id.isin(test_study_ids)]\n```\nThese would be your labels. The paths to each study would be your features, given that you are taking some image as input."
  },
  "source": "meta"
}