{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"#ファイルパスのチェック(Check filepath)\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-09T06:35:17.713059Z","iopub.execute_input":"2022-07-09T06:35:17.713544Z","iopub.status.idle":"2022-07-09T06:35:17.754041Z","shell.execute_reply.started":"2022-07-09T06:35:17.713434Z","shell.execute_reply":"2022-07-09T06:35:17.752754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 0.コンペ概要(About competition) ","metadata":{}},{"cell_type":"markdown","source":"【日本語】<br>\nレストランでの食事やコンサートのチケット購入など、現代の生活では日々の買い物にクレジットカードの利便性が欠かせません。クレジットカードがあれば、多額の現金を持ち歩く必要がなく、また、購入した商品の全額を前払いして、長期にわたって支払うことができます。しかし、カード発行会社は、私たちが請求した金額をきちんと返済してくれることをどうやって確認するのでしょうか？この問題は複雑で、多くの解決策がありますが、このコンペティションでは、さらに多くの改善策を検討します。<br>\n\n貸し倒れ予測は、消費者金融ビジネスのリスク管理の中心的存在です。貸し倒れを予測することで、貸し出しの決定を最適化し、より良い顧客体験と健全なビジネス経済を実現することができます。現在のモデルは、リスク管理を支援するために存在しています。しかし、現在使用されているモデルを凌駕する、より優れたモデルを作成することは可能です。<br>\n\nアメリカン・エキスプレスは、世界的に統合された決済企業です。世界最大の決済カード発行会社である同社は、生活を豊かにし、ビジネスの成功を築くための商品、洞察、体験へのアクセスを顧客に提供しています。<br>\n\nこのコンペティションでは、機械学習のスキルを応用して、クレジット・デフォルトを予測します。具体的には、産業界規模のデータセットを活用し、現在の生産モデルに挑戦する機械学習モデルを構築していただきます。トレーニング、検証、テストの各データセットには、時系列行動データおよび匿名化された顧客プロファイル情報が含まれます。特徴量の作成から、モデル内でのデータの有機的な利用まで、最も強力なモデルを作るためのあらゆる手法を自由に探求することができます。<br>\n\n成功すれば、クレジットカードの審査が通りやすくなり、カード会員にとってより良い顧客体験の実現に貢献することができます。優れたソリューションは、世界最大のクレジットカード発行会社が使用しているクレジットデフォルト予測モデルに挑戦し、賞金やアメリカン・エキスプレスとの面接の機会、そしてやりがいのある新しいキャリアを獲得する可能性があります。<br>\n\nwww.DeepL.com/Translator（無料版）で翻訳しました。\n\n【ENG】<br>\nWhether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.<br>\n\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.<br>\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.<br>\n\nIn this competition, you’ll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.<br>\n\nIf successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.<br>","metadata":{}},{"cell_type":"markdown","source":"# 1. ライブラリインポート(Import library)","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. データの読み込み(Read a dataset)","metadata":{}},{"cell_type":"markdown","source":"## 2.1. データの概形確認(Check data shape)","metadata":{}},{"cell_type":"code","source":"#test_data.csv(33.82 GiB) train_data.csv(16.39 GiB) とサイズが大きいため、データサイズの削減が必要(This dataset is too large size data so, it is need to reduce data size)\n#一旦サンプルで読み込んで中身を確認\ndf1 = pd.read_csv('../input/amex-default-prediction/train_data.csv', nrows=100_000)\n# display(df1)\ndf2 = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\n# display(df2)\ndf = df1.merge(df2,on='customer_ID',how='left')\ndisplay(df)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T07:39:05.675196Z","iopub.execute_input":"2022-07-09T07:39:05.675714Z","iopub.status.idle":"2022-07-09T07:39:11.194618Z","shell.execute_reply.started":"2022-07-09T07:39:05.675681Z","shell.execute_reply":"2022-07-09T07:39:11.193468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#カラムの確認(check column)\ndf.columns.values","metadata":{"execution":{"iopub.status.busy":"2022-07-09T07:39:11.196541Z","iopub.execute_input":"2022-07-09T07:39:11.196914Z","iopub.status.idle":"2022-07-09T07:39:11.204843Z","shell.execute_reply.started":"2022-07-09T07:39:11.196884Z","shell.execute_reply":"2022-07-09T07:39:11.20348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.<br>\nこのコンペティションの目的は、毎月の顧客プロファイルに基づいて、ある顧客が将来クレジットカードの残高を返さない確率を予測することです。対象の二項変数は、最新のクレジットカード明細書から18ヶ月間のパフォーマンスウィンドウを観察して計算され、顧客が最新の明細書の日付から120日以内に返済額を支払わない場合、デフォルトイベントとみなされます。<br>","metadata":{}},{"cell_type":"markdown","source":"The dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:<br>\nこのデータセットには、各顧客の各明細書日付におけるプロファイルの特徴が集約されています。特徴は匿名化、正規化されており、以下の一般的なカテゴリに分類される。<br>\n\nD_* = Delinquency variables（不履行の変数）<br>\nS_* = Spend variables（支出変数）<br>\nP_* = Payment variables（支払い変数）<br>\nB_* = Balance variables（残高変数）<br>\nR_* = Risk variables（リスク変数）<br>","metadata":{}},{"cell_type":"code","source":"#それぞれの変数に入っている値の確認(Check values of column)\npd.set_option('display.max_columns', 100)\n# 対象となる文字列(Target column)\nfor character in ['D', 'S', 'P', 'B', \"R\"]:\n    df_check = df.copy()\n    # 対象文字列を含む列名を取得(Get column name)\n    column_check = [column for column in df_check.columns if character in column]\n    if 'customer_ID' not in column_check:\n        column_check.insert(0, 'customer_ID')\n    df_check = df_check[column_check].set_index(\"customer_ID\")\n    display(df_check)\n    print(df_check.info())\n    display(df_check.describe())","metadata":{"execution":{"iopub.status.busy":"2022-07-09T07:49:06.857025Z","iopub.execute_input":"2022-07-09T07:49:06.857492Z","iopub.status.idle":"2022-07-09T07:49:09.212321Z","shell.execute_reply.started":"2022-07-09T07:49:06.857457Z","shell.execute_reply":"2022-07-09T07:49:09.211077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2. 圧縮版データセットの読み込み(Read a compresssed dataset)","metadata":{}},{"cell_type":"markdown","source":"事前に圧縮版データセットを自身のフォルダに追加する「+ Add data」（Add a compressed dataset to your folder）<br>\nご参考（Reference）AMEX-Feather-Dataset(https://www.kaggle.com/datasets/munumbutt/amexfeather)<br>","metadata":{}},{"cell_type":"code","source":"#ファイルパスのチェック(Check filepath)\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T08:05:49.298335Z","iopub.execute_input":"2022-07-09T08:05:49.298798Z","iopub.status.idle":"2022-07-09T08:05:49.307171Z","shell.execute_reply.started":"2022-07-09T08:05:49.298762Z","shell.execute_reply":"2022-07-09T08:05:49.306108Z"},"trusted":true},"execution_count":null,"outputs":[]}]}