{
  "id": 357318,
  "title": "Weekly learning plan to understand the competition and make a final submission",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/357318",
  "author_name": "",
  "post_date": "2022-10-04T00:17:55.999419Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This is my plan for this competition and I will update everything that I found in each week</p>\n<p>If you find something that answers any of my question, I will be happy to added here.</p>\n<p>Oct-2 to Oct-8: </p>\n<h2>Understand the concepts of Online Machine Learning</h2>\n<ul>\n<li><p>There are good examples by <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> in this discussion <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356584\" target=\"_blank\">post</a></p></li>\n<li><p>My initial  guess is that it might be related with MLOps so streaming processing tools can be helpful</p></li>\n<li><p>Scikit learn I think can also be use based on the post I share, that can be something to take into account</p></li>\n</ul>\n<p>I will try to answer:</p>\n<p>What's Online Machine Learning and what other names can be known?</p>\n<p>Definition:</p>\n<p>The usual approach for batch learning is the way we already understand statistics, we get a csv file from our local computer, we get 1,000 rows and compute the mean and standard deviation of the 1,000 observations and called a day. </p>\n<p>Online learning, process the data row by row, so in this case there is no mean and standard deviation to compute since you have access to the dataset one row at the time, we can calculate running statistics instead, but the idea is that you don't get all the data in one step, you get data one row at the time. </p>\n<pre><code>for xi, yi in zip(X, y):\n    # model learns here\n    pass\n</code></pre>\n<blockquote>\n  <p>The name Online Learning has several problem when we try to look for information. The <a href=\"https://riverml.xyz/0.13.0/examples/batch-to-online/\" target=\"_blank\">river</a> documentation provide alternative to this name such as <strong>Incremental learning</strong> and <strong>stream learning</strong> which are better terms to search. </p>\n  <p>For stream learning there are several tools that we can use such as river,  <a href=\"https://www.tensorflow.org/io/tutorials/kafka\" target=\"_blank\">kafka and tensorflow</a>, and <a href=\"https://spark.apache.org/mllib/\" target=\"_blank\">Spark's MLlib</a> this last one is a work around, and not stream per se, but is something that it worth to learn. </p>\n</blockquote>\n<p>Example of use <strong>River</strong></p>\n<pre><code>from river import compat\nfrom river import compose\n\n# We define a Pipeline, exactly like we did earlier for sklearn \nmodel = compose.Pipeline(\n    ('scale', preprocessing.StandardScaler()),\n    ('log_reg', linear_model.LogisticRegression())\n)\n\n# We make the Pipeline compatible with sklearn\nmodel = compat.convert_river_to_sklearn(model)\n\n# We compute the CV scores using the same CV scheme and the same scoring\nscores = model_selection.cross_val_score(model, X, y, scoring=scorer, cv=cv)\n\n# Display the average score and it's standard deviation\nprint(f'ROC AUC: {scores.mean():.3f} (± {scores.std():.3f})')\n</code></pre>\n<p>What resources are best to understand this concept in a rush?</p>\n<blockquote>\n  <p>Just look at the <a href=\"https://riverml.xyz/0.13.0/examples/batch-to-online/\" target=\"_blank\">documentation</a>, and use this template to create your model</p>\n</blockquote>\n<p>What other kaggle competition use this type of tool to came up with a solution?</p>\n<h2>Check the discussion forum</h2>\n<p>I already found some ideas:</p>\n<ul>\n<li><p>Decrease the size of the dataset (I use a combination of multiple sources)</p></li>\n<li><p>Search for resources that allow to understand Online Machine Learning</p></li>\n<li><p>Watch a video of the game (Extremely helpful) credits to <a href=\"https://www.kaggle.com/mahdeemushfiquekamal\" target=\"_blank\">@mahdeemushfiquekamal</a> in this <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356553\" target=\"_blank\">post</a></p></li>\n</ul>\n<h2>Understand the game</h2>\n<ul>\n<li>Watching a video might be enough, I already have some experience with football (soccer) dataset so maybe I will bring some stuff from that</li>\n</ul>\n<h2>Check NFL notebooks</h2>\n<p>I will check the competition, maybe check YouTube videos and see if anyone already did an analysis of that competition</p>\n<ul>\n<li>Search on YouTube for analysis of the competition</li>\n</ul>\n<blockquote>\n  <p>I found this <a href=\"https://www.youtube.com/watch?v=_Srv0bKmfjY\" target=\"_blank\">video</a> by the winners of the NFL competition<br>\n  Key points from that competition so far:</p>\n  <ol>\n  <li><p>They didn't have domain knowledge, so they understand the basics of the game and use the data science knowledge</p></li>\n  <li><p>Spent time understanding the features, what each one represents and the objective</p></li>\n  <li><p>The goal on that competition that use tracking data (similar to this competition) was to predict the speed of the player after receiving the ball, from that I am guessing it was a ReLu function since negative speed is not possible and there is no (theoretical) limit to the speed of the player. In our problem the goal is to predict the probability of scoring, so this is bounded by 0 and 1, which might implied sigmoid function. </p></li>\n  <li><p>You need to test several things, one out of 5 attemps will probably fail, you need to try a lot.</p></li>\n  <li><p>The winning solution of NFL use a neural network, which is probably a good idea to try here, but I need to check more on that and see what can be use. </p></li>\n  <li><p>Simplify the game might help a lot, the just came up with something that allow to only use the runner and ignore the teammates, that simplies about what the game is doing. </p></li>\n  <li><p>I need to figure out how to handle missing values, I probably will test neural net without using online learning. </p>\n  <p><strong>Next thing to do:</strong></p></li>\n  </ol>\n  <ul>\n  <li></li>\n  <li>Check this <a href=\"https://www.nfl.com/news/next-gen-stats-intro-to-expected-rushing-yards\" target=\"_blank\">post</a> that gives the general approach of the problem</li>\n  <li>Check the kaggle discussion for the winning <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/119400\" target=\"_blank\">solution</a></li>\n  </ul>\n</blockquote>\n<ul>\n<li>Check notebooks for that competition</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/code/jccampos/nfl-2020-winner-solution-the-zoo\" target=\"_blank\">https://www.kaggle.com/code/jccampos/nfl-2020-winner-solution-the-zoo</a></p>\n<ul>\n<li>Check discussion forums for EDA</li>\n</ul>\n<h2>Create the plan for next week</h2>\n<p>Explore some features of your data and develop a plan to follow for the next three weeks</p>\n<h2>Rough plan:</h2>\n<ul>\n<li><p>First week understand the competition and the tools that need to be use</p></li>\n<li><p>Second week perform EDA and find metrics that allow to feature engineering the dataset</p></li>\n<li><p>Third week came up with a baseline model for the dataset</p></li>\n<li><p>Fourth week tune the parameters and came up with the best solution with the resources you have</p></li>\n<li><p>Fifth week look at the winning solution, replicate and try to get insights</p></li>\n<li><p>Sixth week share if you had to start over what would you change. </p></li>\n</ul>",
  "messages": [
    {
      "id": "1970222",
      "postDate": "10/04/2022 00:17:56",
      "content": "<p>This is my plan for this competition and I will update everything that I found in each week</p>\n<p>If you find something that answers any of my question, I will be happy to added here.</p>\n<p>Oct-2 to Oct-8: </p>\n<h2>Understand the concepts of Online Machine Learning</h2>\n<ul>\n<li><p>There are good examples by <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> in this discussion <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356584\" target=\"_blank\">post</a></p></li>\n<li><p>My initial  guess is that it might be related with MLOps so streaming processing tools can be helpful</p></li>\n<li><p>Scikit learn I think can also be use based on the post I share, that can be something to take into account</p></li>\n</ul>\n<p>I will try to answer:</p>\n<p>What's Online Machine Learning and what other names can be known?</p>\n<p>Definition:</p>\n<p>The usual approach for batch learning is the way we already understand statistics, we get a csv file from our local computer, we get 1,000 rows and compute the mean and standard deviation of the 1,000 observations and called a day. </p>\n<p>Online learning, process the data row by row, so in this case there is no mean and standard deviation to compute since you have access to the dataset one row at the time, we can calculate running statistics instead, but the idea is that you don't get all the data in one step, you get data one row at the time. </p>\n<pre><code>for xi, yi in zip(X, y):\n    # model learns here\n    pass\n</code></pre>\n<blockquote>\n  <p>The name Online Learning has several problem when we try to look for information. The <a href=\"https://riverml.xyz/0.13.0/examples/batch-to-online/\" target=\"_blank\">river</a> documentation provide alternative to this name such as <strong>Incremental learning</strong> and <strong>stream learning</strong> which are better terms to search. </p>\n  <p>For stream learning there are several tools that we can use such as river,  <a href=\"https://www.tensorflow.org/io/tutorials/kafka\" target=\"_blank\">kafka and tensorflow</a>, and <a href=\"https://spark.apache.org/mllib/\" target=\"_blank\">Spark's MLlib</a> this last one is a work around, and not stream per se, but is something that it worth to learn. </p>\n</blockquote>\n<p>Example of use <strong>River</strong></p>\n<pre><code>from river import compat\nfrom river import compose\n\n# We define a Pipeline, exactly like we did earlier for sklearn \nmodel = compose.Pipeline(\n    ('scale', preprocessing.StandardScaler()),\n    ('log_reg', linear_model.LogisticRegression())\n)\n\n# We make the Pipeline compatible with sklearn\nmodel = compat.convert_river_to_sklearn(model)\n\n# We compute the CV scores using the same CV scheme and the same scoring\nscores = model_selection.cross_val_score(model, X, y, scoring=scorer, cv=cv)\n\n# Display the average score and it's standard deviation\nprint(f'ROC AUC: {scores.mean():.3f} (± {scores.std():.3f})')\n</code></pre>\n<p>What resources are best to understand this concept in a rush?</p>\n<blockquote>\n  <p>Just look at the <a href=\"https://riverml.xyz/0.13.0/examples/batch-to-online/\" target=\"_blank\">documentation</a>, and use this template to create your model</p>\n</blockquote>\n<p>What other kaggle competition use this type of tool to came up with a solution?</p>\n<h2>Check the discussion forum</h2>\n<p>I already found some ideas:</p>\n<ul>\n<li><p>Decrease the size of the dataset (I use a combination of multiple sources)</p></li>\n<li><p>Search for resources that allow to understand Online Machine Learning</p></li>\n<li><p>Watch a video of the game (Extremely helpful) credits to <a href=\"https://www.kaggle.com/mahdeemushfiquekamal\" target=\"_blank\">@mahdeemushfiquekamal</a> in this <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356553\" target=\"_blank\">post</a></p></li>\n</ul>\n<h2>Understand the game</h2>\n<ul>\n<li>Watching a video might be enough, I already have some experience with football (soccer) dataset so maybe I will bring some stuff from that</li>\n</ul>\n<h2>Check NFL notebooks</h2>\n<p>I will check the competition, maybe check YouTube videos and see if anyone already did an analysis of that competition</p>\n<ul>\n<li>Search on YouTube for analysis of the competition</li>\n</ul>\n<blockquote>\n  <p>I found this <a href=\"https://www.youtube.com/watch?v=_Srv0bKmfjY\" target=\"_blank\">video</a> by the winners of the NFL competition<br>\n  Key points from that competition so far:</p>\n  <ol>\n  <li><p>They didn't have domain knowledge, so they understand the basics of the game and use the data science knowledge</p></li>\n  <li><p>Spent time understanding the features, what each one represents and the objective</p></li>\n  <li><p>The goal on that competition that use tracking data (similar to this competition) was to predict the speed of the player after receiving the ball, from that I am guessing it was a ReLu function since negative speed is not possible and there is no (theoretical) limit to the speed of the player. In our problem the goal is to predict the probability of scoring, so this is bounded by 0 and 1, which might implied sigmoid function. </p></li>\n  <li><p>You need to test several things, one out of 5 attemps will probably fail, you need to try a lot.</p></li>\n  <li><p>The winning solution of NFL use a neural network, which is probably a good idea to try here, but I need to check more on that and see what can be use. </p></li>\n  <li><p>Simplify the game might help a lot, the just came up with something that allow to only use the runner and ignore the teammates, that simplies about what the game is doing. </p></li>\n  <li><p>I need to figure out how to handle missing values, I probably will test neural net without using online learning. </p>\n  <p><strong>Next thing to do:</strong></p></li>\n  </ol>\n  <ul>\n  <li></li>\n  <li>Check this <a href=\"https://www.nfl.com/news/next-gen-stats-intro-to-expected-rushing-yards\" target=\"_blank\">post</a> that gives the general approach of the problem</li>\n  <li>Check the kaggle discussion for the winning <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/119400\" target=\"_blank\">solution</a></li>\n  </ul>\n</blockquote>\n<ul>\n<li>Check notebooks for that competition</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/code/jccampos/nfl-2020-winner-solution-the-zoo\" target=\"_blank\">https://www.kaggle.com/code/jccampos/nfl-2020-winner-solution-the-zoo</a></p>\n<ul>\n<li>Check discussion forums for EDA</li>\n</ul>\n<h2>Create the plan for next week</h2>\n<p>Explore some features of your data and develop a plan to follow for the next three weeks</p>\n<h2>Rough plan:</h2>\n<ul>\n<li><p>First week understand the competition and the tools that need to be use</p></li>\n<li><p>Second week perform EDA and find metrics that allow to feature engineering the dataset</p></li>\n<li><p>Third week came up with a baseline model for the dataset</p></li>\n<li><p>Fourth week tune the parameters and came up with the best solution with the resources you have</p></li>\n<li><p>Fifth week look at the winning solution, replicate and try to get insights</p></li>\n<li><p>Sixth week share if you had to start over what would you change. </p></li>\n</ul>",
      "rawMarkdown": "This is my plan for this competition and I will update everything that I found in each week\n\nIf you find something that answers any of my question, I will be happy to added here.\n\nOct-2 to Oct-8: \n\n## Understand the concepts of Online Machine Learning\n\n- There are good examples by @ravi20076 in this discussion [post](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356584)\n\n- My initial  guess is that it might be related with MLOps so streaming processing tools can be helpful\n\n- Scikit learn I think can also be use based on the post I share, that can be something to take into account\n\nI will try to answer:\n\nWhat's Online Machine Learning and what other names can be known?\n\nDefinition:\n\nThe usual approach for batch learning is the way we already understand statistics, we get a csv file from our local computer, we get 1,000 rows and compute the mean and standard deviation of the 1,000 observations and called a day. \n\nOnline learning, process the data row by row, so in this case there is no mean and standard deviation to compute since you have access to the dataset one row at the time, we can calculate running statistics instead, but the idea is that you don't get all the data in one step, you get data one row at the time. \n\n```python\nfor xi, yi in zip(X, y):\n\t# model learns here\n\tpass\n```\n\n> The name Online Learning has several problem when we try to look for information. The [river](https://riverml.xyz/0.13.0/examples/batch-to-online/) documentation provide alternative to this name such as **Incremental learning** and **stream learning** which are better terms to search. \n> \n> For stream learning there are several tools that we can use such as river,  [kafka and tensorflow](https://www.tensorflow.org/io/tutorials/kafka), and [Spark's MLlib](https://spark.apache.org/mllib/) this last one is a work around, and not stream per se, but is something that it worth to learn. \n\nExample of use **River**\n\n```python\nfrom river import compat\nfrom river import compose\n\n# We define a Pipeline, exactly like we did earlier for sklearn \nmodel = compose.Pipeline(\n    ('scale', preprocessing.StandardScaler()),\n    ('log_reg', linear_model.LogisticRegression())\n)\n\n# We make the Pipeline compatible with sklearn\nmodel = compat.convert_river_to_sklearn(model)\n\n# We compute the CV scores using the same CV scheme and the same scoring\nscores = model_selection.cross_val_score(model, X, y, scoring=scorer, cv=cv)\n\n# Display the average score and it's standard deviation\nprint(f'ROC AUC: {scores.mean():.3f} (± {scores.std():.3f})')\n\n```\n\nWhat resources are best to understand this concept in a rush?\n\n> Just look at the [documentation](https://riverml.xyz/0.13.0/examples/batch-to-online/), and use this template to create your model\n\nWhat other kaggle competition use this type of tool to came up with a solution?\n\n## Check the discussion forum\n\nI already found some ideas:\n\n- Decrease the size of the dataset (I use a combination of multiple sources)\n\n- Search for resources that allow to understand Online Machine Learning\n\n- Watch a video of the game (Extremely helpful) credits to @mahdeemushfiquekamal in this [post](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356553)\n\n## Understand the game\n\n- Watching a video might be enough, I already have some experience with football (soccer) dataset so maybe I will bring some stuff from that\n\n## Check NFL notebooks \n\nI will check the competition, maybe check YouTube videos and see if anyone already did an analysis of that competition\n\n- Search on YouTube for analysis of the competition\n\n> I found this [video](https://www.youtube.com/watch?v=_Srv0bKmfjY) by the winners of the NFL competition\nKey points from that competition so far:\n\n> 1. They didn't have domain knowledge, so they understand the basics of the game and use the data science knowledge\n\n> 2. Spent time understanding the features, what each one represents and the objective\n\n> 3. The goal on that competition that use tracking data (similar to this competition) was to predict the speed of the player after receiving the ball, from that I am guessing it was a ReLu function since negative speed is not possible and there is no (theoretical) limit to the speed of the player. In our problem the goal is to predict the probability of scoring, so this is bounded by 0 and 1, which might implied sigmoid function. \n\n> 4. You need to test several things, one out of 5 attemps will probably fail, you need to try a lot.\n\n> 5. The winning solution of NFL use a neural network, which is probably a good idea to try here, but I need to check more on that and see what can be use. \n\n> 6. Simplify the game might help a lot, the just came up with something that allow to only use the runner and ignore the teammates, that simplies about what the game is doing. \n\n> 7. I need to figure out how to handle missing values, I probably will test neural net without using online learning. \n\n\n>  **Next thing to do:**\n> - ~~Listen the whole video ~~\n> - Check this [post](https://www.nfl.com/news/next-gen-stats-intro-to-expected-rushing-yards) that gives the general approach of the problem\n> - Check the kaggle discussion for the winning [solution](https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/119400)\n\n- Check notebooks for that competition\n\nhttps://www.kaggle.com/code/jccampos/nfl-2020-winner-solution-the-zoo\n\n- Check discussion forums for EDA\n\n## Create the plan for next week\n\nExplore some features of your data and develop a plan to follow for the next three weeks\n\n## Rough plan:\n\n- First week understand the competition and the tools that need to be use\n\n- Second week perform EDA and find metrics that allow to feature engineering the dataset\n\n- Third week came up with a baseline model for the dataset\n\n- Fourth week tune the parameters and came up with the best solution with the resources you have\n\n- Fifth week look at the winning solution, replicate and try to get insights\n\n- Sixth week share if you had to start over what would you change.",
      "votes": null
    },
    {
      "id": "1970233",
      "postDate": "10/04/2022 00:38:11",
      "content": "<p>Thanks for mentioning my post. Best wishes, <a href=\"https://www.kaggle.com/pastorsoto\" target=\"_blank\">@pastorsoto</a></p>",
      "rawMarkdown": "Thanks for mentioning my post. Best wishes, @pastorsoto",
      "votes": null
    },
    {
      "id": "1970274",
      "postDate": "10/04/2022 01:47:14",
      "content": "<p>Thanks for the mention, happy to help anytime <a href=\"https://www.kaggle.com/pastorsoto\" target=\"_blank\">@pastorsoto</a>!</p>",
      "rawMarkdown": "Thanks for the mention, happy to help anytime @pastorsoto!",
      "votes": null
    },
    {
      "id": "1971485",
      "postDate": "10/04/2022 16:33:40",
      "content": "<p>Thanks to you!</p>",
      "rawMarkdown": "Thanks to you!",
      "votes": null
    },
    {
      "id": "1971486",
      "postDate": "10/04/2022 16:34:04",
      "content": "<p>Thanks to you! That video was very helpful</p>",
      "rawMarkdown": "Thanks to you! That video was very helpful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1970233,
      "author_name": "mahdeemushfiquekamal",
      "author_url": "",
      "post_date": "10/04/2022 00:38:11",
      "content": "<p>Thanks for mentioning my post. Best wishes, <a href=\"https://www.kaggle.com/pastorsoto\" target=\"_blank\">@pastorsoto</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1971486,
          "author_name": "pastorsoto",
          "author_url": "",
          "post_date": "10/04/2022 16:34:04",
          "content": "<p>Thanks to you! That video was very helpful</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1970274,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/04/2022 01:47:14",
      "content": "<p>Thanks for the mention, happy to help anytime <a href=\"https://www.kaggle.com/pastorsoto\" target=\"_blank\">@pastorsoto</a>!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1971485,
          "author_name": "pastorsoto",
          "author_url": "",
          "post_date": "10/04/2022 16:33:40",
          "content": "<p>Thanks to you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1970222": "This is my plan for this competition and I will update everything that I found in each week\n\nIf you find something that answers any of my question, I will be happy to added here.\n\nOct-2 to Oct-8: \n\n## Understand the concepts of Online Machine Learning\n\n- There are good examples by @ravi20076 in this discussion [post](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356584)\n\n- My initial  guess is that it might be related with MLOps so streaming processing tools can be helpful\n\n- Scikit learn I think can also be use based on the post I share, that can be something to take into account\n\nI will try to answer:\n\nWhat's Online Machine Learning and what other names can be known?\n\nDefinition:\n\nThe usual approach for batch learning is the way we already understand statistics, we get a csv file from our local computer, we get 1,000 rows and compute the mean and standard deviation of the 1,000 observations and called a day. \n\nOnline learning, process the data row by row, so in this case there is no mean and standard deviation to compute since you have access to the dataset one row at the time, we can calculate running statistics instead, but the idea is that you don't get all the data in one step, you get data one row at the time. \n\n```python\nfor xi, yi in zip(X, y):\n\t# model learns here\n\tpass\n```\n\n> The name Online Learning has several problem when we try to look for information. The [river](https://riverml.xyz/0.13.0/examples/batch-to-online/) documentation provide alternative to this name such as **Incremental learning** and **stream learning** which are better terms to search. \n> \n> For stream learning there are several tools that we can use such as river,  [kafka and tensorflow](https://www.tensorflow.org/io/tutorials/kafka), and [Spark's MLlib](https://spark.apache.org/mllib/) this last one is a work around, and not stream per se, but is something that it worth to learn. \n\nExample of use **River**\n\n```python\nfrom river import compat\nfrom river import compose\n\n# We define a Pipeline, exactly like we did earlier for sklearn \nmodel = compose.Pipeline(\n    ('scale', preprocessing.StandardScaler()),\n    ('log_reg', linear_model.LogisticRegression())\n)\n\n# We make the Pipeline compatible with sklearn\nmodel = compat.convert_river_to_sklearn(model)\n\n# We compute the CV scores using the same CV scheme and the same scoring\nscores = model_selection.cross_val_score(model, X, y, scoring=scorer, cv=cv)\n\n# Display the average score and it's standard deviation\nprint(f'ROC AUC: {scores.mean():.3f} (± {scores.std():.3f})')\n\n```\n\nWhat resources are best to understand this concept in a rush?\n\n> Just look at the [documentation](https://riverml.xyz/0.13.0/examples/batch-to-online/), and use this template to create your model\n\nWhat other kaggle competition use this type of tool to came up with a solution?\n\n## Check the discussion forum\n\nI already found some ideas:\n\n- Decrease the size of the dataset (I use a combination of multiple sources)\n\n- Search for resources that allow to understand Online Machine Learning\n\n- Watch a video of the game (Extremely helpful) credits to @mahdeemushfiquekamal in this [post](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356553)\n\n## Understand the game\n\n- Watching a video might be enough, I already have some experience with football (soccer) dataset so maybe I will bring some stuff from that\n\n## Check NFL notebooks \n\nI will check the competition, maybe check YouTube videos and see if anyone already did an analysis of that competition\n\n- Search on YouTube for analysis of the competition\n\n> I found this [video](https://www.youtube.com/watch?v=_Srv0bKmfjY) by the winners of the NFL competition\nKey points from that competition so far:\n\n> 1. They didn't have domain knowledge, so they understand the basics of the game and use the data science knowledge\n\n> 2. Spent time understanding the features, what each one represents and the objective\n\n> 3. The goal on that competition that use tracking data (similar to this competition) was to predict the speed of the player after receiving the ball, from that I am guessing it was a ReLu function since negative speed is not possible and there is no (theoretical) limit to the speed of the player. In our problem the goal is to predict the probability of scoring, so this is bounded by 0 and 1, which might implied sigmoid function. \n\n> 4. You need to test several things, one out of 5 attemps will probably fail, you need to try a lot.\n\n> 5. The winning solution of NFL use a neural network, which is probably a good idea to try here, but I need to check more on that and see what can be use. \n\n> 6. Simplify the game might help a lot, the just came up with something that allow to only use the runner and ignore the teammates, that simplies about what the game is doing. \n\n> 7. I need to figure out how to handle missing values, I probably will test neural net without using online learning. \n\n\n>  **Next thing to do:**\n> - ~~Listen the whole video ~~\n> - Check this [post](https://www.nfl.com/news/next-gen-stats-intro-to-expected-rushing-yards) that gives the general approach of the problem\n> - Check the kaggle discussion for the winning [solution](https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/119400)\n\n- Check notebooks for that competition\n\nhttps://www.kaggle.com/code/jccampos/nfl-2020-winner-solution-the-zoo\n\n- Check discussion forums for EDA\n\n## Create the plan for next week\n\nExplore some features of your data and develop a plan to follow for the next three weeks\n\n## Rough plan:\n\n- First week understand the competition and the tools that need to be use\n\n- Second week perform EDA and find metrics that allow to feature engineering the dataset\n\n- Third week came up with a baseline model for the dataset\n\n- Fourth week tune the parameters and came up with the best solution with the resources you have\n\n- Fifth week look at the winning solution, replicate and try to get insights\n\n- Sixth week share if you had to start over what would you change.",
    "1970233": "Thanks for mentioning my post. Best wishes, @pastorsoto",
    "1970274": "Thanks for the mention, happy to help anytime @pastorsoto!",
    "1971485": "Thanks to you!",
    "1971486": "Thanks to you! That video was very helpful"
  },
  "source": "meta"
}