{
  "id": 277637,
  "title": "Adding Cross Validation using NumPy Arrays",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/277637",
  "author_name": "",
  "post_date": "2021-10-10T13:23:53.748681300Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>In my notebook I’m getting data into numpy arrays  </p>\n<p><a href=\"https://www.kaggle.com/mmellinger66/brain-tumor-kerastuner\" target=\"_blank\">https://www.kaggle.com/mmellinger66/brain-tumor-kerastuner</a></p>\n<p>In cell 4:</p>\n<pre><code>X, y, trainidt = get_all_data_for_train('T1wCE', image_size=32)\n</code></pre>\n<p>trainidt is a mask array of ids. </p>\n<p>For example, if X[100] and X[101] contain images for training id 42 then trainidt[100] and trainidt[101] would both contain a 42. </p>\n<p>So, for CV I’m going to need to extract different training and validation subsets from X,y, trainidt</p>\n<p>How to I get the proper subset for each KFold?</p>\n<p>I created the folds on train_df so each BraTS21ID is associated with a fold</p>\n<pre><code>splits=KFold(n_splits=k,shuffle=True,random_state=42)\n\nfor fold, (train_idx,valid_idx) in enumerate():\n  _trainidt = trainidt[train_idx] # mask??\n  _X = X[ train_idx] # X is numpy!! Won’t work\n  _y =  \n</code></pre>",
  "messages": [
    {
      "id": "1540444",
      "postDate": "10/10/2021 13:23:53",
      "content": "<p>In my notebook I’m getting data into numpy arrays  </p>\n<p><a href=\"https://www.kaggle.com/mmellinger66/brain-tumor-kerastuner\" target=\"_blank\">https://www.kaggle.com/mmellinger66/brain-tumor-kerastuner</a></p>\n<p>In cell 4:</p>\n<pre><code>X, y, trainidt = get_all_data_for_train('T1wCE', image_size=32)\n</code></pre>\n<p>trainidt is a mask array of ids. </p>\n<p>For example, if X[100] and X[101] contain images for training id 42 then trainidt[100] and trainidt[101] would both contain a 42. </p>\n<p>So, for CV I’m going to need to extract different training and validation subsets from X,y, trainidt</p>\n<p>How to I get the proper subset for each KFold?</p>\n<p>I created the folds on train_df so each BraTS21ID is associated with a fold</p>\n<pre><code>splits=KFold(n_splits=k,shuffle=True,random_state=42)\n\nfor fold, (train_idx,valid_idx) in enumerate():\n  _trainidt = trainidt[train_idx] # mask??\n  _X = X[ train_idx] # X is numpy!! Won’t work\n  _y =  \n</code></pre>",
      "rawMarkdown": "In my notebook I’m getting data into numpy arrays  \n\nhttps://www.kaggle.com/mmellinger66/brain-tumor-kerastuner\n\nIn cell 4:\n\n```\nX, y, trainidt = get_all_data_for_train('T1wCE', image_size=32)\n```\n\ntrainidt is a mask array of ids. \n\nFor example, if X[100] and X[101] contain images for training id 42 then trainidt[100] and trainidt[101] would both contain a 42. \n\nSo, for CV I’m going to need to extract different training and validation subsets from X,y, trainidt\n\nHow to I get the proper subset for each KFold?\n\nI created the folds on train_df so each BraTS21ID is associated with a fold\n\n```\nsplits=KFold(n_splits=k,shuffle=True,random_state=42)\n\nfor fold, (train_idx,valid_idx) in enumerate():\n  _trainidt = trainidt[train_idx] # mask??\n  _X = X[ train_idx] # X is numpy!! Won’t work\n  _y =  \n```",
      "votes": null
    },
    {
      "id": "1540867",
      "postDate": "10/11/2021 02:19:08",
      "content": "<p>i did some work on a function to accomplish the task this afternoon.  Didn’t get around to testing it on my model.</p>\n<p>It uses a loop compression to create a mask, which i believe will let me get the proper subset from X, y, trainidt</p>\n<pre><code>def gen_subset(keep_ids:[int], all_ids:[int], features:[int], labels:[int]) -&gt; ([int], [int], [int]):\n    mask = np.array([i in keep_ids for i in all_ids])\n    print(f\"mask={mask}\")\n    x_filtered = features[mask]\n    print(x_filtered)\n    filtered_ids = all_ids[mask]\n    print(filtered_ids)\n    filtered_labels = labels[mask]\n    print(filtered_labels)\n    return filtered_ids, x_filtered, filtered_labels\n</code></pre>\n<p>Some test data:</p>\n<pre><code>ids = np.array([2,2,2,5,5,5,7,7,9,9])\ntrain_fold_ids = np.array([2,9])\nvalid_fold_ids = np.array([5,7])\n\nX = np.array([0,1,2,3,4,5,6,7,8,9])\nlabels = np.array([1,1,1,3,4,5,6,7,0,0])\ntrain_id_mask, X_train, y_train = gen_subset(train_fold_ids, ids, X, labels)\nvalid_id_mask, X_train, y_train = gen_subset(valid_fold_ids, ids, X, labels)\n\nprint(f\"valid_id_mask={valid_id_mask}\")\n</code></pre>\n<p>Output:</p>\n<pre><code>mask=[ True  True  True False False False False False  True  True]\n[0 1 2 8 9]\n[2 2 2 9 9]\n[1 1 1 0 0]\nmask=[False False False  True  True  True  True  True False False]\n[3 4 5 6 7]\n[5 5 5 7 7]\n[3 4 5 6 7]\nvalid_id_mask=[5 5 5 7 7]\n</code></pre>",
      "rawMarkdown": "i did some work on a function to accomplish the task this afternoon.  Didn’t get around to testing it on my model.\n\nIt uses a loop compression to create a mask, which i believe will let me get the proper subset from X, y, trainidt\n\n```python\ndef gen_subset(keep_ids:[int], all_ids:[int], features:[int], labels:[int]) -> ([int], [int], [int]):\n    mask = np.array([i in keep_ids for i in all_ids])\n    print(f\"mask={mask}\")\n    x_filtered = features[mask]\n    print(x_filtered)\n    filtered_ids = all_ids[mask]\n    print(filtered_ids)\n    filtered_labels = labels[mask]\n    print(filtered_labels)\n    return filtered_ids, x_filtered, filtered_labels\n```\n\nSome test data:\n\n```\nids = np.array([2,2,2,5,5,5,7,7,9,9])\ntrain_fold_ids = np.array([2,9])\nvalid_fold_ids = np.array([5,7])\n\nX = np.array([0,1,2,3,4,5,6,7,8,9])\nlabels = np.array([1,1,1,3,4,5,6,7,0,0])\ntrain_id_mask, X_train, y_train = gen_subset(train_fold_ids, ids, X, labels)\nvalid_id_mask, X_train, y_train = gen_subset(valid_fold_ids, ids, X, labels)\n\nprint(f\"valid_id_mask={valid_id_mask}\")\n```\n\nOutput:\n\n```\nmask=[ True  True  True False False False False False  True  True]\n[0 1 2 8 9]\n[2 2 2 9 9]\n[1 1 1 0 0]\nmask=[False False False  True  True  True  True  True False False]\n[3 4 5 6 7]\n[5 5 5 7 7]\n[3 4 5 6 7]\nvalid_id_mask=[5 5 5 7 7]\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1540867,
      "author_name": "mmellinger66",
      "author_url": "",
      "post_date": "10/11/2021 02:19:08",
      "content": "<p>i did some work on a function to accomplish the task this afternoon.  Didn’t get around to testing it on my model.</p>\n<p>It uses a loop compression to create a mask, which i believe will let me get the proper subset from X, y, trainidt</p>\n<pre><code>def gen_subset(keep_ids:[int], all_ids:[int], features:[int], labels:[int]) -&gt; ([int], [int], [int]):\n    mask = np.array([i in keep_ids for i in all_ids])\n    print(f\"mask={mask}\")\n    x_filtered = features[mask]\n    print(x_filtered)\n    filtered_ids = all_ids[mask]\n    print(filtered_ids)\n    filtered_labels = labels[mask]\n    print(filtered_labels)\n    return filtered_ids, x_filtered, filtered_labels\n</code></pre>\n<p>Some test data:</p>\n<pre><code>ids = np.array([2,2,2,5,5,5,7,7,9,9])\ntrain_fold_ids = np.array([2,9])\nvalid_fold_ids = np.array([5,7])\n\nX = np.array([0,1,2,3,4,5,6,7,8,9])\nlabels = np.array([1,1,1,3,4,5,6,7,0,0])\ntrain_id_mask, X_train, y_train = gen_subset(train_fold_ids, ids, X, labels)\nvalid_id_mask, X_train, y_train = gen_subset(valid_fold_ids, ids, X, labels)\n\nprint(f\"valid_id_mask={valid_id_mask}\")\n</code></pre>\n<p>Output:</p>\n<pre><code>mask=[ True  True  True False False False False False  True  True]\n[0 1 2 8 9]\n[2 2 2 9 9]\n[1 1 1 0 0]\nmask=[False False False  True  True  True  True  True False False]\n[3 4 5 6 7]\n[5 5 5 7 7]\n[3 4 5 6 7]\nvalid_id_mask=[5 5 5 7 7]\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1540444": "In my notebook I’m getting data into numpy arrays  \n\nhttps://www.kaggle.com/mmellinger66/brain-tumor-kerastuner\n\nIn cell 4:\n\n```\nX, y, trainidt = get_all_data_for_train('T1wCE', image_size=32)\n```\n\ntrainidt is a mask array of ids. \n\nFor example, if X[100] and X[101] contain images for training id 42 then trainidt[100] and trainidt[101] would both contain a 42. \n\nSo, for CV I’m going to need to extract different training and validation subsets from X,y, trainidt\n\nHow to I get the proper subset for each KFold?\n\nI created the folds on train_df so each BraTS21ID is associated with a fold\n\n```\nsplits=KFold(n_splits=k,shuffle=True,random_state=42)\n\nfor fold, (train_idx,valid_idx) in enumerate():\n  _trainidt = trainidt[train_idx] # mask??\n  _X = X[ train_idx] # X is numpy!! Won’t work\n  _y =  \n```",
    "1540867": "i did some work on a function to accomplish the task this afternoon.  Didn’t get around to testing it on my model.\n\nIt uses a loop compression to create a mask, which i believe will let me get the proper subset from X, y, trainidt\n\n```python\ndef gen_subset(keep_ids:[int], all_ids:[int], features:[int], labels:[int]) -> ([int], [int], [int]):\n    mask = np.array([i in keep_ids for i in all_ids])\n    print(f\"mask={mask}\")\n    x_filtered = features[mask]\n    print(x_filtered)\n    filtered_ids = all_ids[mask]\n    print(filtered_ids)\n    filtered_labels = labels[mask]\n    print(filtered_labels)\n    return filtered_ids, x_filtered, filtered_labels\n```\n\nSome test data:\n\n```\nids = np.array([2,2,2,5,5,5,7,7,9,9])\ntrain_fold_ids = np.array([2,9])\nvalid_fold_ids = np.array([5,7])\n\nX = np.array([0,1,2,3,4,5,6,7,8,9])\nlabels = np.array([1,1,1,3,4,5,6,7,0,0])\ntrain_id_mask, X_train, y_train = gen_subset(train_fold_ids, ids, X, labels)\nvalid_id_mask, X_train, y_train = gen_subset(valid_fold_ids, ids, X, labels)\n\nprint(f\"valid_id_mask={valid_id_mask}\")\n```\n\nOutput:\n\n```\nmask=[ True  True  True False False False False False  True  True]\n[0 1 2 8 9]\n[2 2 2 9 9]\n[1 1 1 0 0]\nmask=[False False False  True  True  True  True  True False False]\n[3 4 5 6 7]\n[5 5 5 7 7]\n[3 4 5 6 7]\nvalid_id_mask=[5 5 5 7 7]\n```"
  },
  "source": "meta"
}