{
  "id": 245668,
  "title": "Multi-Index vs. Dummy Variables",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/245668",
  "author_name": "",
  "post_date": "2021-06-11T20:03:04.649357600Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I noticed something interesting when looking at the model in the default notebook: the model input is based on a multi-index as opposed to typical dummy variables usually seen to determine categorical variables.</p>\n<p>More specifically, the y_train input has a columnar index like this:<br>\n{'target1': [player1, player2, …, playern], …, 'target4': [player1, player2, …, playern]}.</p>\n<p>and the X_train columnar index is composed of tuples:<br>\n[(runsScored, player1), …, (runsScored, playern), …, (totalBases, player1), …, (totalBases, playern), sin(1, freq=A-DEC), …]</p>\n<p>It seems this is a nice way of allowing the X_train and y_train dataframes to only have one row observation per date, and one doesn't need to duplicate data by basically having copies of the fourier columns for every individual player.</p>\n<p>A few questions that follow:</p>\n<ol>\n<li><p>Is there a name for this approach to make it easily search-engine-able?</p></li>\n<li><p>How would we go about predicting on an out-of-sample player?<br>\nWould we want to retrain, replicate the data across all inputs and then average the output targets, some other method?</p></li>\n<li><p>Does the model know to only associate the columns marked player# with the target1-4 columns for each player or does the training method essentially set the gradients of each of the nodes in the network at such a weight that helps to isolate its value against others?  If so, would retraining the model on another row of date data expect to significantly alter the player-specific weights?  What about removing or adding more players: would we expect the weights associated with the relationship between player# X to player# y to change drastically?</p></li>\n</ol>\n<p>If there is a link to a blog or other online source that helps explain this method, feel free to point me in that direction and avoid answering my somewhat lengthy questions entirely.  This is a neat method; I just haven't seen it before!</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "1345740",
      "postDate": "06/11/2021 20:03:04",
      "content": "<p>I noticed something interesting when looking at the model in the default notebook: the model input is based on a multi-index as opposed to typical dummy variables usually seen to determine categorical variables.</p>\n<p>More specifically, the y_train input has a columnar index like this:<br>\n{'target1': [player1, player2, …, playern], …, 'target4': [player1, player2, …, playern]}.</p>\n<p>and the X_train columnar index is composed of tuples:<br>\n[(runsScored, player1), …, (runsScored, playern), …, (totalBases, player1), …, (totalBases, playern), sin(1, freq=A-DEC), …]</p>\n<p>It seems this is a nice way of allowing the X_train and y_train dataframes to only have one row observation per date, and one doesn't need to duplicate data by basically having copies of the fourier columns for every individual player.</p>\n<p>A few questions that follow:</p>\n<ol>\n<li><p>Is there a name for this approach to make it easily search-engine-able?</p></li>\n<li><p>How would we go about predicting on an out-of-sample player?<br>\nWould we want to retrain, replicate the data across all inputs and then average the output targets, some other method?</p></li>\n<li><p>Does the model know to only associate the columns marked player# with the target1-4 columns for each player or does the training method essentially set the gradients of each of the nodes in the network at such a weight that helps to isolate its value against others?  If so, would retraining the model on another row of date data expect to significantly alter the player-specific weights?  What about removing or adding more players: would we expect the weights associated with the relationship between player# X to player# y to change drastically?</p></li>\n</ol>\n<p>If there is a link to a blog or other online source that helps explain this method, feel free to point me in that direction and avoid answering my somewhat lengthy questions entirely.  This is a neat method; I just haven't seen it before!</p>\n<p>Thank you!</p>",
      "rawMarkdown": "I noticed something interesting when looking at the model in the default notebook: the model input is based on a multi-index as opposed to typical dummy variables usually seen to determine categorical variables.\n\nMore specifically, the y_train input has a columnar index like this:\n{'target1': [player1, player2, ..., playern], ..., 'target4': [player1, player2, ..., playern]}.\n\nand the X_train columnar index is composed of tuples:\n[(runsScored, player1), ..., (runsScored, playern), ..., (totalBases, player1), ..., (totalBases, playern), sin(1, freq=A-DEC), ...]\n\nIt seems this is a nice way of allowing the X_train and y_train dataframes to only have one row observation per date, and one doesn't need to duplicate data by basically having copies of the fourier columns for every individual player.\n\nA few questions that follow:\n1. Is there a name for this approach to make it easily search-engine-able?\n\n2. How would we go about predicting on an out-of-sample player?\nWould we want to retrain, replicate the data across all inputs and then average the output targets, some other method?\n\n3. Does the model know to only associate the columns marked player# with the target1-4 columns for each player or does the training method essentially set the gradients of each of the nodes in the network at such a weight that helps to isolate its value against others?  If so, would retraining the model on another row of date data expect to significantly alter the player-specific weights?  What about removing or adding more players: would we expect the weights associated with the relationship between player# X to player# y to change drastically?\n\nIf there is a link to a blog or other online source that helps explain this method, feel free to point me in that direction and avoid answering my somewhat lengthy questions entirely.  This is a neat method; I just haven't seen it before!\n\nThank you!",
      "votes": null
    },
    {
      "id": "1346002",
      "postDate": "06/12/2021 03:18:22",
      "content": "<p>Nicely written. Upvoted. Do take a look at my notebooks and give feedback.</p>",
      "rawMarkdown": "Nicely written. Upvoted. Do take a look at my notebooks and give feedback.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1346002,
      "author_name": "arnab132",
      "author_url": "",
      "post_date": "06/12/2021 03:18:22",
      "content": "<p>Nicely written. Upvoted. Do take a look at my notebooks and give feedback.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1345740": "I noticed something interesting when looking at the model in the default notebook: the model input is based on a multi-index as opposed to typical dummy variables usually seen to determine categorical variables.\n\nMore specifically, the y_train input has a columnar index like this:\n{'target1': [player1, player2, ..., playern], ..., 'target4': [player1, player2, ..., playern]}.\n\nand the X_train columnar index is composed of tuples:\n[(runsScored, player1), ..., (runsScored, playern), ..., (totalBases, player1), ..., (totalBases, playern), sin(1, freq=A-DEC), ...]\n\nIt seems this is a nice way of allowing the X_train and y_train dataframes to only have one row observation per date, and one doesn't need to duplicate data by basically having copies of the fourier columns for every individual player.\n\nA few questions that follow:\n1. Is there a name for this approach to make it easily search-engine-able?\n\n2. How would we go about predicting on an out-of-sample player?\nWould we want to retrain, replicate the data across all inputs and then average the output targets, some other method?\n\n3. Does the model know to only associate the columns marked player# with the target1-4 columns for each player or does the training method essentially set the gradients of each of the nodes in the network at such a weight that helps to isolate its value against others?  If so, would retraining the model on another row of date data expect to significantly alter the player-specific weights?  What about removing or adding more players: would we expect the weights associated with the relationship between player# X to player# y to change drastically?\n\nIf there is a link to a blog or other online source that helps explain this method, feel free to point me in that direction and avoid answering my somewhat lengthy questions entirely.  This is a neat method; I just haven't seen it before!\n\nThank you!",
    "1346002": "Nicely written. Upvoted. Do take a look at my notebooks and give feedback."
  },
  "source": "meta"
}