{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"},{"sourceId":201377683,"sourceType":"kernelVersion"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style = \"font-weight: 900; font-family : Calibri; font-size : 40px\">FOREWORD</h1>\n\nThis kernel is a starter EDA kernel for the Jane Street 2024 competition. <br>\nWe have 79 features here, from feature_00 - feature_78 and a few join-key columns like date_id, time_id and stock symbol_id <br>\nWe also have 9 targets - responder_0 - responder_8 and we need to predict the next day's value for responder_6 only. <br>\n\nThis is a time series regression problem os a non-equi-spaced nature. Conventional methods are unlikely to work here in my opinion. <br>\n\nIn this kernel, we delve into the data and try and unearth possible patterns that could be used to build features and even drop useless ones <br> Let's see how this develops over the next days <br>\n\n\n- Version 1 of this notebook contains plots for feature cluster 2 <br>\n- Version 2 of this notebook contains plots for feature cluster 0 <br>\n- Version 3 of this notebook contains plots for feature cluster 4 <br>","metadata":{}},{"cell_type":"markdown","source":"<h1 style = \"font-weight: 900; font-family : Calibri; font-size : 40px\">IMPORTS</h1>","metadata":{}},{"cell_type":"code","source":"%%time \n\n!pip install polars[gpu]==1.9.0  -q  --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/polars\n!pip install lightgbm==4.5.0     -q  --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/packages\n!pip install scikit-learn==1.5.2 -q  --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/packages\n\nexec(\n    open(\"/kaggle/input/janestreet2024-imports-v1/myimports.py\", \"r\"\n        ).read()\n)\n\nprint()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style = \"font-weight: 900; font-family : Calibri; font-size : 40px; font-weight : 900\">CONFIGURATION<h1>","metadata":{}},{"cell_type":"code","source":"%%time \n\neda_strt_dt = 1000\neda_end_dt  = 1400\nversion_nb  = \"V1_1\"\n\nip_path     = f\"/kaggle/input/jane-street-real-time-market-data-forecasting\"\nop_path     = f\"/kaggle/working\"\nstate       = 42\n\ntarget      = \"responder_6\"\n\n\n# Plot axes styler:-\nmyaxesstyle = \\\n{'axes.facecolor'   : 'white',\n 'axes.edgecolor'   : 'black',\n 'axes.grid'        : False,\n 'axes.axisbelow'   : True,\n 'axes.labelcolor'  : '.15',\n 'figure.facecolor' : 'white',\n 'grid.color'       : 'white',\n 'grid.linestyle'   : '-',\n 'patch.edgecolor'  : 'white',\n 'figure.facecolor' : 'white',\n}\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\ntarget_cols = \\\n[\"responder_0\", \"responder_1\", \"responder_2\", \"responder_3\", \"responder_4\",\n \"responder_5\", \"responder_6\" , \"responder_7\", \"responder_8\",\n]\n\nkey_cols = [\"date_id\", \"time_id\", \"symbol_id\"]\n\nstrt_cols = \\\nkey_cols + \\\n[\n'weight',\n'feature_00', 'feature_01', 'feature_02', 'feature_03', 'feature_04',\n'feature_05', 'feature_06', 'feature_07', 'feature_08', 'feature_09', 'feature_10',\n'feature_11', 'feature_12', 'feature_13', 'feature_14', 'feature_15', 'feature_16',\n'feature_17', 'feature_18', 'feature_19', 'feature_20', 'feature_21', 'feature_22',\n'feature_23', 'feature_24', 'feature_25', 'feature_26', 'feature_27', 'feature_28',\n'feature_29', 'feature_30', 'feature_31', 'feature_32', 'feature_33', 'feature_34',\n'feature_35', 'feature_36', 'feature_37', 'feature_38', 'feature_39', 'feature_40',\n'feature_41', 'feature_42', 'feature_43', 'feature_44', 'feature_45', 'feature_46',\n'feature_47', 'feature_48', 'feature_49', 'feature_50', 'feature_51', 'feature_52',\n'feature_53', 'feature_54', 'feature_55', 'feature_56', 'feature_57', 'feature_58',\n'feature_59', 'feature_60', 'feature_61', 'feature_62', 'feature_63', 'feature_64',\n'feature_65', 'feature_66', 'feature_67', 'feature_68', 'feature_69', 'feature_70',\n'feature_71', 'feature_72', 'feature_73', 'feature_74', 'feature_75', 'feature_76',\n'feature_77', 'feature_78',\n]","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style = \"font-weight: 900; font-family : Calibri; font-size : 40px; font-weight : 900\">DATA LOAD</h1>\n\nWe load a sample dataset here, using all columns from the input data over sampled date_ids. We retain the continuity of dates for time series meaningfulness <br>","metadata":{}},{"cell_type":"code","source":"%%time \n\ntrain = \\\n(pl.scan_parquet(os.path.join(ip_path, \"train.parquet\")).\n filter(pl.col(\"date_id\").is_between(eda_strt_dt, eda_end_dt)).\n select(pl.col(strt_cols))\n)\n\ntrain_ftretgt = \\\n(pl.scan_parquet(os.path.join(ip_path, \"train.parquet\")).\n filter(pl.col(\"date_id\").is_between(eda_strt_dt, eda_end_dt)).\n drop(\"partition_id\", strict = False)\n)\n\nfull_train = \\\npl.scan_parquet(os.path.join(ip_path, \"train.parquet\"))","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style = \"font-weight: 900; font-family : Calibri; font-size : 40px; font-weight : 900\">EDA</h1>","metadata":{}},{"cell_type":"markdown","source":"<h2 style = \"font-weight: 900; font-family : Calibri; font-weight : 900; font-size : 32px\">DESCRIPTIONS </h2>\n\nLet's delve into the features and look for potentially interesting ones with basic descriptions and unique values","metadata":{}},{"cell_type":"code","source":"%%time \n\ndf = \\\npl.concat(\n    [train.select(pl.col(strt_cols[4:])).describe(),\n     train.select(pl.col(strt_cols[4:]).n_unique()).\n     collect().\n     insert_column(0, pl.Series(\"statistic\", [\"n_unique\"]))\n    ],\n    how = \"vertical_relaxed\",\n)\n    \ndf.write_csv(f\"Descriptions_{version_nb}.csv\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\ndf_1 = df.to_pandas().set_index(\"statistic\").transpose()\ndf_1[\"ratio_unq_cnt\"]   = df_1[\"n_unique\"] / df_1[\"count\"]\ndf_1[\"ratio_nulls_cnt\"] = df_1[\"null_count\"] / df_1[\"count\"]\ndf_1[\"range\"]    = df_1[\"max\"] - df_1[\"min\"]\ndf_1[\"dir_skew\"] = np.where(df_1[\"mean\"] > df_1[\"50%\"], \"right\", \"left\")\n\nfmt1 = \"color : red; fontweight: bold; background-color : white;  border: dashed maroon 1.2px\"\nfmt2 = \"color : blue; fontweight: bold; background-color : white; border: dashed brown 1.2px\"\n\ndisplay(\n    (df_1.style.\n     applymap(\n         lambda x: fmt1 if x > 0.50 else fmt2,\n         subset = [\"ratio_unq_cnt\"]\n     ).\n     applymap(\n         lambda x: fmt1 if x > 0.0001 else fmt2,\n         subset = [\"ratio_nulls_cnt\"]\n     ).\n     applymap(\n         lambda x: fmt1 if x < 1000 else fmt2,\n         subset = [\"n_unique\"]\n     ). \n     applymap(\n         lambda x: fmt1 if x == \"right\" else fmt2,\n         subset = [\"dir_skew\"]\n     ).  \n     applymap(\n         lambda x: fmt1 if x > 100 else fmt2,\n         subset = [\"range\"]\n     ).       \n     format(formatter = {\"count\"      : \"{:,.0f}\",\n                         \"n_unique\"   : \"{:,.0f}\",\n                         \"null_count\" : \"{:,.0f}\",\n                        },\n            precision = 4,\n           ).\n     set_caption(f\"Basic descriptions\")\n    )\n)\n\ndel df, df_1\ncollect();","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color: white; \n            border: 2px dashed maroon; \n            padding: 15px; \n            margin: 10px; \n            border-radius: 8px;\n            box-shadow: 3px 3px 5px #888888;\">\n<h3 style=\"color: black; margin-top: 0; font-weight: 900; font-family : Calibri\"> INFERENCES</h3>\nFeatures 9, 10, 11 appear to be categorical features, these need to be studied separately <br>\nFeature 61 is completely useless and can be dropped <br>\nFeatures from 20-31 have low unique counts but they are floats, again worthy of reconsideration <br>\nFeatures like feature_06 have a very high range and could be considered for further exploration\n</div>","metadata":{}},{"cell_type":"markdown","source":"<h2 style = \"font-weight: 900; font-family : Calibri; font-weight : 900; font-size : 32px\">SYMBOL-IDs </h2>\n\nLet's delve into the start and end dates of symbol ids and check for their availability through dates","metadata":{}},{"cell_type":"code","source":"%%time \n\ntry: \n    del df\nexcept:\n    pass\n\ndf = \\\n(full_train.\n select([\"date_id\", \"symbol_id\", \"time_id\"]).\n group_by([\"date_id\", \"symbol_id\"]).\n agg(pl.col([\"time_id\"]).count().alias(\"counts\")).\n collect().\n to_pandas().\n pivot_table(\n     index      = \"date_id\",\n     columns    = \"symbol_id\",\n     values     = \"counts\",\n     aggfunc    = \"sum\",\n     fill_value = 0,\n )\n)\n\ndf[\"nb_symbols\"] = df.clip(0, 1).sum(axis=1)\n\ndisplay(\n    df[[\"nb_symbols\"]].\n    sort_values([\"nb_symbols\"], ascending = True).\n    head(10).\n    transpose().\n    style.\n    set_caption(\"Lowest 10 symbols occurrances\")\n)\n\ndisplay(\n    df[[\"nb_symbols\"]].\n    sort_values([\"nb_symbols\"], ascending = True).\n    tail(10).\n    transpose().\n    style.\n    set_caption(\"All symbols occurances\")\n)\n\nprint()\ncollect();","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\nwith np.printoptions(linewidth = 150):\n    PrintColor(f\"---> Dates with less than 10 symbols\")\n    pprint(\n        np.array(df.loc[df.nb_symbols < 10].index)\n    )\n    \n    PrintColor(f\"\\n\\n---> Dates with symbols between 11 and 20\")\n    pprint(\n        np.array(df.loc[df.nb_symbols.between(11, 20)].index)\n    )\n    \n    PrintColor(f\"\\n\\n---> Dates with symbols between 21 and 30\")\n    pprint(\n        np.array(df.loc[df.nb_symbols.between(21, 30)].index)\n    )    \n    \n    PrintColor(f\"\\n\\n---> Dates with symbols between 31 and 37\")\n    pprint(\n        np.array(df.loc[df.nb_symbols.between(31, 37)].index)\n    ) \n       \n    PrintColor(f\"\\n\\n---> Dates with all symbols\")\n    pprint(\n        np.array(df.loc[df.nb_symbols.between(38, 40)].index)\n    )        ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\nwith sns.axes_style(\"whitegrid\"):\n    fig, ax = plt.subplots(1,1, figsize = (25,8))\n\n    df.iloc[:, 0: -1].clip(0,1).sum().plot.bar(ax = ax, color = \"tab:blue\")\n    ax.set_title(f\"Symbol occurrances by date\\n\", \n                 **{\"fontsize\" : 14, \"color\" : \"maroon\", \"fontweight\": \"bold\"}\n                )\n    ax.axhline(y = 500, **{\"linewidth\" : 2.5, \"color\" : \"maroon\", \"linestyle\" : \"dashed\"})\n    ax.axhline(y = 1000, **{\"linewidth\" : 2.5, \"color\" : \"maroon\", \"linestyle\" : \"dashed\"})\n    ax.axhline(y = 1600, **{\"linewidth\" : 2.5, \"color\" : \"maroon\", \"linestyle\" : \"dashed\"})\n    ax.set_xticks(range(0, 39, 1), labels = range(0, 39, 1), rotation = 0)\n    ax.set_yticks(range(0, 1701, 100), labels = range(0, 1701, 100), rotation = 0)  \n    \n    plt.tight_layout()\n    plt.show()\n    \ndel df\ncollect()\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color: white; \n            border: 2px dashed maroon; \n            padding: 15px; \n            margin: 10px; \n            border-radius: 8px;\n            box-shadow: 3px 3px 5px #888888;\">\n<h3 style=\"color: black; margin-top: 0; font-weight: 900; font-family : Calibri\"> INFERENCES</h3>\nSymbols 18, 4, 6, 24, 31, 32 are relatively rare <br>\nSymbols 1, 7 19 occur almost on all dates <br>\nMost of the recent dates have most of the symbols <br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<h2 style = \"font-weight: 900; font-family : Calibri; font-weight : 900; font-size : 32px\">WEIGHT </h2>\n\nLet's delve into the weights column and check it across dates and symbols","metadata":{}},{"cell_type":"code","source":"%%time \n\nwith sns.axes_style(myaxesstyle):\n    df = \\\n    pl.sql(\n        \"\"\"\n        SELECT date_id, symbol_id, weight, count(date_id) as num_rcrds\n        FROM full_train\n        group by date_id, symbol_id, weight \n        order by date_id, symbol_id, weight \n        \"\"\"\n    )\n    all_colors = sns.color_palette(\"icefire\", n_colors = 100,)\n    \n    fig, axes = \\\n    plt.subplots(13, 3, figsize = (33, 102), \n                 gridspec_kw = {\"hspace\" : 0.2, \"wspace\" : 0.2}\n                )\n    \n    for i in range(39):\n        ax = axes[i//3, i % 3]\n        j  = np.random.randint(0, 100,)\n        \n        (df.\n         filter(pl.col(\"symbol_id\").eq(i)).\n         select([\"date_id\", \"weight\"]).\n         collect().\n         to_pandas().\n         set_index(\"date_id\").\n         plot.line(ax = ax, color = all_colors[j])\n        )\n        ax.set_title(f\"Symbol{i}\", **{\"fontweight\" : 900, \"color\" : \"brown\"})\n        ax.set(xlabel = \"\")\n        ax.set_xticks(\n            range(0, 1701, 100), labels = range(0, 1701, 100), rotation = 90\n        )\n        \n    plt.suptitle(\n        f\"Analysis between weight by symbol and date\", \n        y = 0.89, \n        **{\"fontweight\" : 900, \"color\" : \"black\"}\n    )\n    plt.tight_layout()\n    plt.show()\n    \ncollect();     ","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \n\ndef make_ftre_plot(\n    ftre: str, target = \"responder_6\", myaxesstyle: dict = myaxesstyle\n):\n    \"This function makes a line-plot between the 2 provided features/ columns across symbols\"\n            \n    with sns.axes_style(myaxesstyle):\n        fig, axes = \\\n        plt.subplots(13, 3, figsize = (33, 102), \n                     gridspec_kw = {\"hspace\" : 0.2, \"wspace\" : 0.2}\n                    )\n\n        for i in tqdm(range(0, 39)):\n            \n            mydf = \\\n            (\n            full_train.\n            select(pl.col(key_cols + [ftre, target])).\n            filter(pl.col(\"symbol_id\").eq(i)).\n            group_by(\"date_id\", maintain_order = True).\n            agg(pl.col(ftre).mean().alias(\"mean_ftre\"),\n                pl.col(target).mean().alias(\"mean_tgt\"),\n               ).\n            collect().\n            drop_nulls().\n            to_pandas().\n            set_index(\"date_id\")\n            )\n            corr_ = mydf.dropna().corr()[\"mean_tgt\"][0]\n        \n            ax = axes[i//3, i % 3]\n            mydf.plot.line(ax = ax, color = [\"#66b3ff\", \"#b33c00\"], linewidth = 1.0)\n            \n            ax.set_title(\n                f\"Symbol{i} - correlation = {corr_ :.6f}\", \n                **{\"fontweight\" : 900, \"color\" : \"brown\"}\n            )\n            ax.set(xlabel= \"\")\n\n        plt.suptitle(\n            f\"{ftre.upper()} analysis with target by symbol\", \n            y = 0.89, \n            **{\"fontweight\" : 900, \"color\" : \"black\"}\n        )\n        plt.tight_layout()\n        plt.show()\n        collect();\n","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color: black; margin-top: 0; font-weight: 900; font-family : Calibri; font-size : 20px\"> Let's make feature plots for feature cluster-4 to start with </div>","metadata":{}},{"cell_type":"code","source":"make_ftre_plot(\"feature_18\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_39\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_40\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_41\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_45\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_50\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_51\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_52\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_56\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(\"feature_65\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color: black; margin-top: 0; font-weight: 900; font-family : Calibri; font-size : 20px\"> Let's plot responders against our target to continue the analysis </div>","metadata":{}},{"cell_type":"code","source":"make_ftre_plot(ftre = \"responder_0\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(ftre = \"responder_1\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(ftre = \"responder_2\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(ftre = \"responder_3\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(ftre = \"responder_4\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(ftre = \"responder_5\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(ftre = \"responder_7\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"make_ftre_plot(ftre = \"responder_8\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color: white; \n            border: 2px dashed maroon; \n            padding: 15px; \n            margin: 10px; \n            border-radius: 8px;\n            box-shadow: 3px 3px 5px #888888;\">\n<h3 style=\"color: black; margin-top: 0; font-weight: 900; font-family : Calibri\"> INFERENCES</h3>\nResponder columns seem to hold value - especially responder_4 seems to be well correlated with the target<br>\nResponder_5 seems to be poorly correlated with the target <br>\nEven responder_7 and responder_8 seem to be extremely well correlated with the target <br>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<h3 style=\"color: black; margin-top: 0; font-weight: 900; font-family : Calibri; font-size: 30px\"> TO BE CONTINUED</h3>","metadata":{}}]}