{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Riiid! Answer Correctness Prediction\n## Introduction\n在本次比賽中，您將預測每個學生能夠正確回答的問題。 您將遍歷一系列問題。 做出預測後，您可以繼續進行下一批。\n\n本次比賽與大多數 Kaggle 比賽的不同之處在於：\n* 您只能從 Kaggle Notebooks 提交\n* 您必須使用我們自定義的 **`riiiducation`** Python 模塊。 該模塊的目的是控制信息流，以確保您不會使用未來數據進行預測。 如果您沒有正確使用此模塊，您的代碼可能會失敗。\n\n## 在本入門筆記本中，我們將展示如何使用 **`riiiducation`** 模塊來獲取測試功能並進行預測。\n## TL;DR：端到端使用示例\n```\nimport riiideducation\nenv = riiideducation.make_env()\n\n# 訓練數據照常在比賽數據集中訓練數據照常在比賽數據集中\ntrain_df = pd.read_csv('/kaggle/input/riiideducation/train.csv', low_memory=False)\ntrain_my_model(train_df)\niter_test = env.iter_test()\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df['answered_correctly'] = 0.5\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])```\n請注意，`train_my_model` 是您需要為上述示例工作而編寫的函數。","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true}},{"cell_type":"markdown","source":"## 深入介紹\n首先讓我們導入模塊並創建一個環境。","metadata":{}},{"cell_type":"code","source":"import riiideducation\nimport pandas as pd\n\n# You can only call make_env() once, so don't lose it!\nenv = riiideducation.make_env()","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:26.967421Z","iopub.execute_input":"2022-05-08T16:21:26.968016Z","iopub.status.idle":"2022-05-08T16:21:26.993001Z","shell.execute_reply.started":"2022-05-08T16:21:26.967980Z","shell.execute_reply":"2022-05-08T16:21:26.992084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:27.011306Z","iopub.execute_input":"2022-05-08T16:21:27.011913Z","iopub.status.idle":"2022-05-08T16:21:27.022149Z","shell.execute_reply.started":"2022-05-08T16:21:27.011879Z","shell.execute_reply":"2022-05-08T16:21:27.021431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 訓練數據照常在比賽數據集中\n它比默認設置的內存要大，所以我們將指定更有效的數據類型，現在只加載數據的子集。","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', low_memory=False, nrows=10**5, \n                       dtype={'row_id': 'int64', 'timestamp': 'int64', 'user_id': 'int32', 'content_id': 'int16', 'content_type_id': 'int8',\n                              'task_container_id': 'int16', 'user_answer': 'int8', 'answered_correctly': 'int8', 'prior_question_elapsed_time': 'float32', \n                             'prior_question_had_explanation': 'boolean',\n                             }\n                      )\ntrain_df","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:27.060919Z","iopub.execute_input":"2022-05-08T16:21:27.061611Z","iopub.status.idle":"2022-05-08T16:21:27.346875Z","shell.execute_reply.started":"2022-05-08T16:21:27.061555Z","shell.execute_reply":"2022-05-08T16:21:27.345854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## `iter_test` 函數\n\n循環遍歷測試集中每批問題的生成器。 為方便起見，您可以直接訪問示例測試行，但您的代碼只能通過 API 從真實測試集中獲取行。 調用 **`predict`** 後，您可以繼續下一批。\n\n產量：\n* 雖然有更多批次，並且自上次產量以來成功調用了 `predict`，但會產生一個元組：\n     * `test_df`：具有下一批測試功能的 DataFrame，以及上一批的用戶響應。\n     * `sample_prediction_df`：帶有示例預測的 DataFrame。 旨在填寫並傳遞回“預測”函數。\n* 如果 `predict` 自上次 yield 後未成功調用，則打印錯誤並產生 `None`。","metadata":{}},{"cell_type":"code","source":"# You can only iterate through a result from `env.iter_test()` once\n# so be careful not to lose it once you start iterating.\niter_test = env.iter_test()","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:27.348404Z","iopub.execute_input":"2022-05-08T16:21:27.348869Z","iopub.status.idle":"2022-05-08T16:21:27.352201Z","shell.execute_reply.started":"2022-05-08T16:21:27.348836Z","shell.execute_reply":"2022-05-08T16:21:27.351561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"讓我們獲取第一個測試批次的數據並檢查一下。","metadata":{}},{"cell_type":"code","source":"(test_df, sample_prediction_df) = next(iter_test)\ntest_df","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:27.353379Z","iopub.execute_input":"2022-05-08T16:21:27.353649Z","iopub.status.idle":"2022-05-08T16:21:27.400629Z","shell.execute_reply.started":"2022-05-08T16:21:27.353623Z","shell.execute_reply":"2022-05-08T16:21:27.399635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_prediction_df","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:27.401612Z","iopub.execute_input":"2022-05-08T16:21:27.401871Z","iopub.status.idle":"2022-05-08T16:21:27.412699Z","shell.execute_reply.started":"2022-05-08T16:21:27.401845Z","shell.execute_reply":"2022-05-08T16:21:27.411569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"請注意，如果我們嘗試繼續下一批而不對當前批次進行預測，我們將得到一個錯誤。","metadata":{}},{"cell_type":"code","source":"next(iter_test)","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:27.414735Z","iopub.execute_input":"2022-05-08T16:21:27.415060Z","iopub.status.idle":"2022-05-08T16:21:27.424935Z","shell.execute_reply.started":"2022-05-08T16:21:27.415019Z","shell.execute_reply":"2022-05-08T16:21:27.423783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **`預測`** 函數\n存儲您對當前批次的預測。 期望與您在從 `iter_test` 生成器返回的 `sample_prediction_df` 中看到的格式相同。\n\n參數：\n* `predictions_df`：必須與`sample_prediction_df`具有相同格式的DataFrame。\n\n如果在 `iter_test` 生成器成功迭代後未調用此函數，則會引發異常。","metadata":{}},{"cell_type":"markdown","source":"讓我們使用 `iter_test` 提供的樣本進行虛擬預測。","metadata":{}},{"cell_type":"code","source":"env.predict(sample_prediction_df)","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:27.426257Z","iopub.execute_input":"2022-05-08T16:21:27.426748Z","iopub.status.idle":"2022-05-08T16:21:27.434414Z","shell.execute_reply.started":"2022-05-08T16:21:27.426716Z","shell.execute_reply":"2022-05-08T16:21:27.433394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 主循環\n讓我們遍歷測試集生成器中的所有剩餘批次，並對每個批次進行默認預測。 一旦你到達終點，`iter_test` 生成器將簡單地停止返回值。\n\n在編寫自己的 Notebooks 時，一定要編寫健壯的代碼，盡可能少地對 `iter_test`/`predict` 循環做出假設。 例如，測試集包含以前在訓練中未觀察到的問題 ID。\n\n您可以假設 `sample_prediction_df` 的結構在本次比賽中不會改變。\n\n**不應提交 `test_df` 中的講座行。**","metadata":{}},{"cell_type":"code","source":"for (test_df, sample_prediction_df) in iter_test:\n    test_df['answered_correctly'] = 0.5\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])","metadata":{"execution":{"iopub.status.busy":"2022-05-08T16:21:27.435721Z","iopub.execute_input":"2022-05-08T16:21:27.436025Z","iopub.status.idle":"2022-05-08T16:21:28.074833Z","shell.execute_reply.started":"2022-05-08T16:21:27.435993Z","shell.execute_reply":"2022-05-08T16:21:28.073605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 重新啟動筆記本以再次運行您的代碼\n為了打擊作弊，您只能在每次 Notebook 運行時調用 `make_env` 或迭代 `iter_test` 一次。 然而，當你在你的模型上迭代時，嘗試一些東西是合理的，稍微改變一下模型，然後再試一次。 不幸的是，如果您嘗試簡單地重新運行代碼，甚至刷新瀏覽器頁面，您仍然會在之前運行的同一 Notebook 執行會話上運行，並且 `riideducation` 模塊仍然會拋出錯誤。 要解決此問題，您需要明確重新啟動 Notebook 執行會話，您可以通過**單擊 Notebook Editor 頂部菜單欄中的“Run”->“Restart Session”** 來完成。## Restart the Notebook to run your code again\nIn order to combat cheating, you are only allowed to call `make_env` or iterate through `iter_test` once per Notebook run.  However, while you're iterating on your model it's reasonable to try something out, change the model a bit, and try it again.  Unfortunately, if you try to simply re-run the code, or even refresh the browser page, you'll still be running on the same Notebook execution session you had been running before, and the `riideducation` module will still throw errors.  To get around this, you need to explicitly restart your Notebook execution session, which you can do by **clicking \"Run\"->\"Restart Session\"** in the Notebook Editor's menu bar at the top.","metadata":{"trusted":true}},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}