{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This year's Jane Steer competition has started, and with it a mass of data in the form of anonymized vfeatures. We have a total of 79 features starting with the prefix \"feature\" and 9 features referring to \"responder\". A significant feature (different from the others) is **\"symbol_id\", which refers to a specific, but also anonymized financial instrument**, which is a categorical variable. Financial instruments differ from each other, so it is worth putting forward a **hypothesis whether the type of financial instrument affects the shape of the other variables**. The answer to this question will help us choose modeling strategies: is it worth creating a separate model for each instrument, or is it worth combining new features in feature enigeinnering based on the discovered dependencies. Let's start!","metadata":{"_kg_hide-input":false}},{"cell_type":"markdown","source":"We use Python with additional libraries for analysis. We will start by using data processing libraries (numpy, pandas, polars) and data visualization libraries (matplotlib, seaborn). To increase code readability, we will remove all warnings that pop up during the creation of graphs. Additionally, we will make a very simple switcher called \"debug\" that will allow us to quickly test whether the code works, it will be set to \"true\", in other cases when we want to use all data, we will change the value to \"false\".","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport polars as pl\nimport gc\nimport seaborn as sns\nimport warnings\nimport random\n\nwarnings.filterwarnings('ignore')\n\ndebug = True","metadata":{"execution":{"iopub.status.busy":"2024-11-01T18:35:30.103118Z","iopub.execute_input":"2024-11-01T18:35:30.103664Z","iopub.status.idle":"2024-11-01T18:35:32.127653Z","shell.execute_reply.started":"2024-11-01T18:35:30.103583Z","shell.execute_reply":"2024-11-01T18:35:32.126057Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The data is in .parquet format, divided into 10 smaller parts. We analyze only the training set, because for the test set we have only the file format without specific data, which will be completed only during the model validation stage, so the **entire analysis is based only on the training set**. We combine smaller parts into one set (on non-debug version) and remove unnecessary smaller parts that we will not use, to free up RAM. As symbol_id is a categorical feature written in the form of numbers, for the sake of clarity we change it to the \"string\" format to avoid a case in which it will be mistakenly treated as a numeric feature during calculations for visualization, which will cause us to achieve unintended effects.","metadata":{}},{"cell_type":"code","source":"if debug:\n    df_0 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=0/part-0.parquet')\n    df_1 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=1/part-0.parquet')\n    df_2 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=2/part-0.parquet')\n    df_3 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=3/part-0.parquet')\n    df_4 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=4/part-0.parquet')\n    train = pd.concat([df_0, df_1, df_2, df_3, df_4])\n    del df_0, df_1, df_2, df_3, df_4\n\nelse: \n    df_0 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=0/part-0.parquet')\n    df_1 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=1/part-0.parquet')\n    df_2 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=2/part-0.parquet')\n    df_3 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=3/part-0.parquet')\n    df_4 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=4/part-0.parquet')\n    df_5 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=5/part-0.parquet')\n    df_6 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=6/part-0.parquet')\n    df_7 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=7/part-0.parquet')\n    df_8 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=8/part-0.parquet')\n    df_9 = pd.read_parquet('../input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=9/part-0.parquet')\n    train = pd.concat([df_0, df_1, df_2, df_3, df_4, df_5, df_6, df_7, df_8, df_9])\n    del df_0, df_1, df_2, df_3, df_4, df_5, df_6, df_7, df_8, df_9\n\ngc.collect()\n\ntrain['symbol_id'] = train['symbol_id'].astype(str)\ntrain = train.reset_index()\n#train = train.replace([np.inf, -np.inf], np.nan, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-11-01T18:35:32.129775Z","iopub.execute_input":"2024-11-01T18:35:32.130333Z","iopub.status.idle":"2024-11-01T18:36:45.479175Z","shell.execute_reply.started":"2024-11-01T18:35:32.130290Z","shell.execute_reply":"2024-11-01T18:36:45.476867Z"},"_kg_hide-input":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"To verify the correctness of the above actions, we display the training database to make sure it contains the appropriate amount of data and features that we saw earlier in the preview.","metadata":{}},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2024-11-01T18:36:45.482410Z","iopub.execute_input":"2024-11-01T18:36:45.483059Z","iopub.status.idle":"2024-11-01T18:36:51.296314Z","shell.execute_reply.started":"2024-11-01T18:36:45.482965Z","shell.execute_reply":"2024-11-01T18:36:51.294587Z"},"_kg_hide-input":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Let's get to the point - we want to **check the distributions of features for each financial instrument**. Therefore, we will create a densit plot for each category of the symbol_id column and each quantitative variable. So we create a **separate plot for each feature, where there will be a set of n smaller plots, where n is the number of unique categories of the \"symbol_id\" feature**, i.e. financial instruments. We enclose the whole thing in a function with three arguments: the database (in our case \"train\"), the name of the quantitative feature and the name of the grouping column (in our case \"symbol_id\").","metadata":{}},{"cell_type":"code","source":"def create_plot(df, feature, grouping_feature):\n    title = 'Density by symbol_id for feature: ' + str(feature)\n    g = sns.FacetGrid(df, col=grouping_feature, sharex=False, sharey=False, col_wrap=6)\n    g.map_dataframe(sns.kdeplot, x = feature, fill = True, alpha = 0.75, color=random.choice(['#66c2a5', '#fc8d62', '#8da0cb', '#e78ac3', '#a6d854', '#ffd92f', '#e5c494', '#b3b3b3']))\n    g.set(xlabel=None, ylabel=None)\n    g.fig.subplots_adjust(top=0.92)\n    g.fig.suptitle(title, fontweight='bold', fontsize=25)","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Time to fire up the machine! Using a \"for\" loop for each quantitative feature we use a function where the iterator of the function is the name of that feature. We have almost 100 of them, so the calculations will take some time.","metadata":{}},{"cell_type":"code","source":"list_of_features_part_1 = [i for i in train.columns if i.startswith('feature')]\nlist_of_features_part_2 = [i for i in train.columns if i.startswith('responder')]\n\nlist_of_features = list_of_features_part_1 + list_of_features_part_2 + ['weight']\n\nfor i in train[list_of_features].columns:\n    create_plot(train, i, 'symbol_id')","metadata":{"execution":{"iopub.status.busy":"2024-11-01T18:36:51.314793Z","iopub.execute_input":"2024-11-01T18:36:51.315548Z","iopub.status.idle":"2024-11-01T18:43:36.009271Z","shell.execute_reply.started":"2024-11-01T18:36:51.315469Z","shell.execute_reply":"2024-11-01T18:43:36.007559Z"},"_kg_hide-input":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We can now browse the schedules (which are over 2000), so **analyzing each one separately is too time-consuming in relation to the knowledge acquired in this way, so it is worth focusing on for some category** (one of the financial instruments) if the feature does not behave differently. Enjoy watching! Remember, if we have a better machine it is worth running caluclations on the entire train set, for the needs of this notebook I am running only on the first five of 10 parts.","metadata":{}},{"cell_type":"markdown","source":"**Thanks for reading my notebook!**\n\n**If you have any suggestions for improving the analysis, let me know in the comment!**\n\n**If you appreciate my work in this notebook, give upvote!**\n\n**If you have a moment, I encourage you to see at my other [projects](https://www.kaggle.com/michau96/code).**","metadata":{"_kg_hide-input":false}}]}