{
  "id": 382884,
  "title": "[Silver] 113th - Computationally simple solution",
  "url": "/competitions/otto-recommender-system/writeups/memory-error-silver-113th-computationally-simple-s",
  "author_name": "",
  "post_date": "2023-02-10T10:06:30.823Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p><em>I am currently updating this post and I will share my github repository once I have finished fixing some code</em></p>\n<p>Through this competition I learned a lot, so I thought that sharing my experience might help, in particular because I was able to achieve this score by (almost) using only my laptop, which is a macbook with 16GB of ram. This is my first medal so I am quite happy nevertheless.</p>\n<p>Like almost everyone else I also used the two-step approach:</p>\n<ul>\n<li>Candidate generation</li>\n<li>Re-ranking </li>\n</ul>\n<p>One thing that saved me was a slightly modified version of the function that you can find at the link, which I ran after almost each operation: </p>\n<pre><code>def reduce_memory(df):\n    \"\"\" \n    Credits: https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook\n    Function that iterates through all the columns of a dataframe and modify the data type to reduce memory usage.        \n    \"\"\"\n\n    for col in df.columns:\n        col_type = df[col].dtype\n\n        if (str(col_type)[:3] == 'int') or (str(col_type)[:3] == 'flo') or (str(col_type)[:3] == 'obj'):\n            if col_type != object:\n                c_min = df[col].min()\n                c_max = df[col].max()\n                if str(col_type)[:3] == 'int':\n                    if c_min &gt; np.iinfo(np.int8).min and c_max &lt; np.iinfo(np.int8).max:\n                        df[col] = df[col].astype('int8')\n                    elif c_min &gt; np.iinfo(np.int16).min and c_max &lt; np.iinfo(np.int16).max:\n                        df[col] = df[col].astype('int16')\n                    elif c_min &gt; np.iinfo(np.int32).min and c_max &lt; np.iinfo(np.int32).max:\n                        df[col] = df[col].astype('int32')\n                    elif c_min &gt; np.iinfo(np.int64).min and c_max &lt; np.iinfo(np.int64).max:\n                        df[col] = df[col].astype('int64')  \n                else:\n                    if c_min &gt; np.finfo(np.float16).min and c_max &lt; np.finfo(np.float16).max:\n                        df[col] = df[col].astype('float16')\n                    elif c_min &gt; np.finfo(np.float32).min and c_max &lt; np.finfo(np.float32).max:\n                        df[col] = df[col].astype('float32')\n                    else:\n                        df[col] = df[col].astype('float64')\n            else:\n                df[col] = df[col].astype('category')\n\n    return df\n</code></pre>\n<h2>Candidate Generation</h2>\n<p>I was able to reach the following recall for the validation set</p>\n<ul>\n<li>Clicks: 62.04%</li>\n<li>Carts:  53.89%</li>\n<li>Orders: 73.23%</li>\n</ul>\n<p>through the following approach. The candidates for each session were produced through:</p>\n<ul>\n<li><strong>Co-visitation matrix</strong> with everything in the train data (this is the only thing that I computed through Kaggle notebooks)</li>\n<li><strong>Co-visitation matrix with</strong> everything in the train data with only <strong>positive time distance</strong>, i.e. we link item y to item x only if item y is after item x and within the time frame.</li>\n<li><strong>Co-visitation matrix</strong> with only <strong>carts and orders</strong></li>\n<li><strong>Co-visitation matrix</strong> with only <strong>orders</strong></li>\n<li>Last <strong>previously interacted items</strong></li>\n<li><strong>Word2Vec similar items</strong> to the last one seen in the session</li>\n</ul>\n<p>and to make everything simple, after computing the co-visitation matrices, I split the session into folds through the following code:</p>\n<pre><code>session_array = test['session'].unique()\nrandom.shuffle(session_array)\nsession_split = np.array_split(session_array, N_FOLDS)\ni = 0\nfor session_subset in session_split:\n     if i == 0:\n          fold_df = pd.DataFrame(zip(session_subset, np.zeros(len(session_subset)) ), columns=['session', 'fold'])\n          fold_df['fold'] = fold_df['fold'].astype(int)\n     else:\n          extra = pd.DataFrame(zip(session_subset, np.zeros(len(session_subset)) ), columns=['session', 'fold'])\n          extra['fold'] = extra['fold'].astype(int)\n          extra['fold'] = extra['fold'] + i\n          fold_df = pd.concat([fold_df, extra], axis=0)\n     i = i+1\nfold_df.to_csv('session_fold.csv', index=False)\n</code></pre>\n<p>This allowed me to produce the candidate dataframe separately avoiding memory issues. Therefore for each fold I had the whole set of candidates in separate files. With this I was able to perform undersampling to reduce the # of negative samples to 20 times the # of positives samples, and I saved these dataframes for the training. Having them separately made performing n-fold training and computing validation very easily since I had both sampled and original candidates ready.</p>\n<h2>Features</h2>\n<p>This was probably the part where I was lacking the most due to time constraint on my part and due to the low score that I got, given an high recall rate on the candidate generation part. I produced features regarding:</p>\n<ul>\n<li>user</li>\n<li>item</li>\n<li>user-item interactions</li>\n<li>how the candidate was generated</li>\n</ul>\n<p>and I saved them as csv files locally, in order to be able to combine them with the required dataframe as I needed.</p>\n<h2>Training</h2>\n<p>For training, as I stated before, having the sampled dataframes saved allowed me to combine them easily with the features and also to save them. This made me able to train an XGBoost Ranker locally on my macbook in ~20-30 minutes. I wasn't able to try the LGBM Ranker since it wasn't supported for my current MacOS version (remember to never update to the latest os version haha), while the Catboost ranker has almost the same performances.</p>\n<h2>Inference</h2>\n<p>For inference, to avoid out of memory errors I loaded each non-sampled dataframe by chunk, add features to the chunk and then perform inference and finally saving the predictions. At the end I combined the predictions and took the top 20 ones to submit.  </p>",
  "messages": [
    {
      "id": "2125034",
      "postDate": "02/01/2023 11:28:54",
      "content": "<p><em>I am currently updating this post and I will share my github repository once I have finished fixing some code</em></p>\n<p>Through this competition I learned a lot, so I thought that sharing my experience might help, in particular because I was able to achieve this score by (almost) using only my laptop, which is a macbook with 16GB of ram. This is my first medal so I am quite happy nevertheless.</p>\n<p>Like almost everyone else I also used the two-step approach:</p>\n<ul>\n<li>Candidate generation</li>\n<li>Re-ranking </li>\n</ul>\n<p>One thing that saved me was a slightly modified version of the function that you can find at the link, which I ran after almost each operation: </p>\n<pre><code>def reduce_memory(df):\n    \"\"\" \n    Credits: https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook\n    Function that iterates through all the columns of a dataframe and modify the data type to reduce memory usage.        \n    \"\"\"\n\n    for col in df.columns:\n        col_type = df[col].dtype\n\n        if (str(col_type)[:3] == 'int') or (str(col_type)[:3] == 'flo') or (str(col_type)[:3] == 'obj'):\n            if col_type != object:\n                c_min = df[col].min()\n                c_max = df[col].max()\n                if str(col_type)[:3] == 'int':\n                    if c_min &gt; np.iinfo(np.int8).min and c_max &lt; np.iinfo(np.int8).max:\n                        df[col] = df[col].astype('int8')\n                    elif c_min &gt; np.iinfo(np.int16).min and c_max &lt; np.iinfo(np.int16).max:\n                        df[col] = df[col].astype('int16')\n                    elif c_min &gt; np.iinfo(np.int32).min and c_max &lt; np.iinfo(np.int32).max:\n                        df[col] = df[col].astype('int32')\n                    elif c_min &gt; np.iinfo(np.int64).min and c_max &lt; np.iinfo(np.int64).max:\n                        df[col] = df[col].astype('int64')  \n                else:\n                    if c_min &gt; np.finfo(np.float16).min and c_max &lt; np.finfo(np.float16).max:\n                        df[col] = df[col].astype('float16')\n                    elif c_min &gt; np.finfo(np.float32).min and c_max &lt; np.finfo(np.float32).max:\n                        df[col] = df[col].astype('float32')\n                    else:\n                        df[col] = df[col].astype('float64')\n            else:\n                df[col] = df[col].astype('category')\n\n    return df\n</code></pre>\n<h2>Candidate Generation</h2>\n<p>I was able to reach the following recall for the validation set</p>\n<ul>\n<li>Clicks: 62.04%</li>\n<li>Carts:  53.89%</li>\n<li>Orders: 73.23%</li>\n</ul>\n<p>through the following approach. The candidates for each session were produced through:</p>\n<ul>\n<li><strong>Co-visitation matrix</strong> with everything in the train data (this is the only thing that I computed through Kaggle notebooks)</li>\n<li><strong>Co-visitation matrix with</strong> everything in the train data with only <strong>positive time distance</strong>, i.e. we link item y to item x only if item y is after item x and within the time frame.</li>\n<li><strong>Co-visitation matrix</strong> with only <strong>carts and orders</strong></li>\n<li><strong>Co-visitation matrix</strong> with only <strong>orders</strong></li>\n<li>Last <strong>previously interacted items</strong></li>\n<li><strong>Word2Vec similar items</strong> to the last one seen in the session</li>\n</ul>\n<p>and to make everything simple, after computing the co-visitation matrices, I split the session into folds through the following code:</p>\n<pre><code>session_array = test['session'].unique()\nrandom.shuffle(session_array)\nsession_split = np.array_split(session_array, N_FOLDS)\ni = 0\nfor session_subset in session_split:\n     if i == 0:\n          fold_df = pd.DataFrame(zip(session_subset, np.zeros(len(session_subset)) ), columns=['session', 'fold'])\n          fold_df['fold'] = fold_df['fold'].astype(int)\n     else:\n          extra = pd.DataFrame(zip(session_subset, np.zeros(len(session_subset)) ), columns=['session', 'fold'])\n          extra['fold'] = extra['fold'].astype(int)\n          extra['fold'] = extra['fold'] + i\n          fold_df = pd.concat([fold_df, extra], axis=0)\n     i = i+1\nfold_df.to_csv('session_fold.csv', index=False)\n</code></pre>\n<p>This allowed me to produce the candidate dataframe separately avoiding memory issues. Therefore for each fold I had the whole set of candidates in separate files. With this I was able to perform undersampling to reduce the # of negative samples to 20 times the # of positives samples, and I saved these dataframes for the training. Having them separately made performing n-fold training and computing validation very easily since I had both sampled and original candidates ready.</p>\n<h2>Features</h2>\n<p>This was probably the part where I was lacking the most due to time constraint on my part and due to the low score that I got, given an high recall rate on the candidate generation part. I produced features regarding:</p>\n<ul>\n<li>user</li>\n<li>item</li>\n<li>user-item interactions</li>\n<li>how the candidate was generated</li>\n</ul>\n<p>and I saved them as csv files locally, in order to be able to combine them with the required dataframe as I needed.</p>\n<h2>Training</h2>\n<p>For training, as I stated before, having the sampled dataframes saved allowed me to combine them easily with the features and also to save them. This made me able to train an XGBoost Ranker locally on my macbook in ~20-30 minutes. I wasn't able to try the LGBM Ranker since it wasn't supported for my current MacOS version (remember to never update to the latest os version haha), while the Catboost ranker has almost the same performances.</p>\n<h2>Inference</h2>\n<p>For inference, to avoid out of memory errors I loaded each non-sampled dataframe by chunk, add features to the chunk and then perform inference and finally saving the predictions. At the end I combined the predictions and took the top 20 ones to submit.  </p>",
      "rawMarkdown": "*I am currently updating this post and I will share my github repository once I have finished fixing some code*\n\nThrough this competition I learned a lot, so I thought that sharing my experience might help, in particular because I was able to achieve this score by (almost) using only my laptop, which is a macbook with 16GB of ram. This is my first medal so I am quite happy nevertheless.\n\nLike almost everyone else I also used the two-step approach:\n - Candidate generation\n - Re-ranking \n\nOne thing that saved me was a slightly modified version of the function that you can find at the link, which I ran after almost each operation: \n```\ndef reduce_memory(df):\n    \"\"\" \n    Credits: https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook\n    Function that iterates through all the columns of a dataframe and modify the data type to reduce memory usage.        \n    \"\"\"\n    \n    for col in df.columns:\n        col_type = df[col].dtype\n        \n        if (str(col_type)[:3] == 'int') or (str(col_type)[:3] == 'flo') or (str(col_type)[:3] == 'obj'):\n            if col_type != object:\n                c_min = df[col].min()\n                c_max = df[col].max()\n                if str(col_type)[:3] == 'int':\n                    if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                        df[col] = df[col].astype('int8')\n                    elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                        df[col] = df[col].astype('int16')\n                    elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                        df[col] = df[col].astype('int32')\n                    elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                        df[col] = df[col].astype('int64')  \n                else:\n                    if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                        df[col] = df[col].astype('float16')\n                    elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                        df[col] = df[col].astype('float32')\n                    else:\n                        df[col] = df[col].astype('float64')\n            else:\n                df[col] = df[col].astype('category')\n\n    return df\n```\n\n## Candidate Generation\nI was able to reach the following recall for the validation set\n - Clicks: 62.04%\n - Carts:  53.89%\n - Orders: 73.23%\n\nthrough the following approach. The candidates for each session were produced through:\n - **Co-visitation matrix** with everything in the train data (this is the only thing that I computed through Kaggle notebooks)\n - **Co-visitation matrix with** everything in the train data with only **positive time distance**, i.e. we link item y to item x only if item y is after item x and within the time frame.\n - **Co-visitation matrix** with only **carts and orders**\n - **Co-visitation matrix** with only **orders**\n - Last **previously interacted items**\n - **Word2Vec similar items** to the last one seen in the session\n\nand to make everything simple, after computing the co-visitation matrices, I split the session into folds through the following code:\n\n```\nsession_array = test['session'].unique()\nrandom.shuffle(session_array)\nsession_split = np.array_split(session_array, N_FOLDS)\ni = 0\nfor session_subset in session_split:\n     if i == 0:\n          fold_df = pd.DataFrame(zip(session_subset, np.zeros(len(session_subset)) ), columns=['session', 'fold'])\n          fold_df['fold'] = fold_df['fold'].astype(int)\n     else:\n          extra = pd.DataFrame(zip(session_subset, np.zeros(len(session_subset)) ), columns=['session', 'fold'])\n          extra['fold'] = extra['fold'].astype(int)\n          extra['fold'] = extra['fold'] + i\n          fold_df = pd.concat([fold_df, extra], axis=0)\n     i = i+1\nfold_df.to_csv('session_fold.csv', index=False)\n```\n\nThis allowed me to produce the candidate dataframe separately avoiding memory issues. Therefore for each fold I had the whole set of candidates in separate files. With this I was able to perform undersampling to reduce the # of negative samples to 20 times the # of positives samples, and I saved these dataframes for the training. Having them separately made performing n-fold training and computing validation very easily since I had both sampled and original candidates ready.\n\n## Features\nThis was probably the part where I was lacking the most due to time constraint on my part and due to the low score that I got, given an high recall rate on the candidate generation part. I produced features regarding:\n - user\n - item\n - user-item interactions\n - how the candidate was generated\n\nand I saved them as csv files locally, in order to be able to combine them with the required dataframe as I needed.\n\n## Training\nFor training, as I stated before, having the sampled dataframes saved allowed me to combine them easily with the features and also to save them. This made me able to train an XGBoost Ranker locally on my macbook in ~20-30 minutes. I wasn't able to try the LGBM Ranker since it wasn't supported for my current MacOS version (remember to never update to the latest os version haha), while the Catboost ranker has almost the same performances.\n\n## Inference\nFor inference, to avoid out of memory errors I loaded each non-sampled dataframe by chunk, add features to the chunk and then perform inference and finally saving the predictions. At the end I combined the predictions and took the top 20 ones to submit.",
      "votes": null
    },
    {
      "id": "2125035",
      "postDate": "02/01/2023 11:30:19",
      "content": "<p>Congrats! No worries after the leaderboard is finalized you will probably get silver ;)</p>",
      "rawMarkdown": "Congrats! No worries after the leaderboard is finalized you will probably get silver ;)",
      "votes": null
    },
    {
      "id": "2125126",
      "postDate": "02/01/2023 12:49:35",
      "content": "<p>Thank you! I hope so, having that massive drop in ranks and landing outside silver wasn't funny at all…</p>\n<p>Also congrats on your score and on your silver medal!</p>",
      "rawMarkdown": "Thank you! I hope so, having that massive drop in ranks and landing outside silver wasn't funny at all...\n\nAlso congrats on your score and on your silver medal!",
      "votes": null
    },
    {
      "id": "2125358",
      "postDate": "02/01/2023 15:59:14",
      "content": "<blockquote>\n  <p>Word2Vec similar items to the last one seen in the session<br>\n  It is worth trying also min/mean/max to several last aids in session. I guess this alone could give you enough points to stay in silver.</p>\n</blockquote>",
      "rawMarkdown": ">Word2Vec similar items to the last one seen in the session\nIt is worth trying also min/mean/max to several last aids in session. I guess this alone could give you enough points to stay in silver.",
      "votes": null
    },
    {
      "id": "2125454",
      "postDate": "02/01/2023 17:04:01",
      "content": "<p>This idea was on my to do list, together with trying different co-visitation matrices but unfortunately I didn't have time. </p>\n<p>Congrats on your medal!</p>",
      "rawMarkdown": "This idea was on my to do list, together with trying different co-visitation matrices but unfortunately I didn't have time. \n\nCongrats on your medal!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2125035,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "02/01/2023 11:30:19",
      "content": "<p>Congrats! No worries after the leaderboard is finalized you will probably get silver ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125126,
          "author_name": "mattiabrazzale",
          "author_url": "",
          "post_date": "02/01/2023 12:49:35",
          "content": "<p>Thank you! I hope so, having that massive drop in ranks and landing outside silver wasn't funny at all…</p>\n<p>Also congrats on your score and on your silver medal!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125358,
      "author_name": "artemfedorov",
      "author_url": "",
      "post_date": "02/01/2023 15:59:14",
      "content": "<blockquote>\n  <p>Word2Vec similar items to the last one seen in the session<br>\n  It is worth trying also min/mean/max to several last aids in session. I guess this alone could give you enough points to stay in silver.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 2125454,
          "author_name": "mattiabrazzale",
          "author_url": "",
          "post_date": "02/01/2023 17:04:01",
          "content": "<p>This idea was on my to do list, together with trying different co-visitation matrices but unfortunately I didn't have time. </p>\n<p>Congrats on your medal!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2125034": "*I am currently updating this post and I will share my github repository once I have finished fixing some code*\n\nThrough this competition I learned a lot, so I thought that sharing my experience might help, in particular because I was able to achieve this score by (almost) using only my laptop, which is a macbook with 16GB of ram. This is my first medal so I am quite happy nevertheless.\n\nLike almost everyone else I also used the two-step approach:\n - Candidate generation\n - Re-ranking \n\nOne thing that saved me was a slightly modified version of the function that you can find at the link, which I ran after almost each operation: \n```\ndef reduce_memory(df):\n    \"\"\" \n    Credits: https://www.kaggle.com/code/gemartin/load-data-reduce-memory-usage/notebook\n    Function that iterates through all the columns of a dataframe and modify the data type to reduce memory usage.        \n    \"\"\"\n    \n    for col in df.columns:\n        col_type = df[col].dtype\n        \n        if (str(col_type)[:3] == 'int') or (str(col_type)[:3] == 'flo') or (str(col_type)[:3] == 'obj'):\n            if col_type != object:\n                c_min = df[col].min()\n                c_max = df[col].max()\n                if str(col_type)[:3] == 'int':\n                    if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                        df[col] = df[col].astype('int8')\n                    elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                        df[col] = df[col].astype('int16')\n                    elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                        df[col] = df[col].astype('int32')\n                    elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                        df[col] = df[col].astype('int64')  \n                else:\n                    if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                        df[col] = df[col].astype('float16')\n                    elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                        df[col] = df[col].astype('float32')\n                    else:\n                        df[col] = df[col].astype('float64')\n            else:\n                df[col] = df[col].astype('category')\n\n    return df\n```\n\n## Candidate Generation\nI was able to reach the following recall for the validation set\n - Clicks: 62.04%\n - Carts:  53.89%\n - Orders: 73.23%\n\nthrough the following approach. The candidates for each session were produced through:\n - **Co-visitation matrix** with everything in the train data (this is the only thing that I computed through Kaggle notebooks)\n - **Co-visitation matrix with** everything in the train data with only **positive time distance**, i.e. we link item y to item x only if item y is after item x and within the time frame.\n - **Co-visitation matrix** with only **carts and orders**\n - **Co-visitation matrix** with only **orders**\n - Last **previously interacted items**\n - **Word2Vec similar items** to the last one seen in the session\n\nand to make everything simple, after computing the co-visitation matrices, I split the session into folds through the following code:\n\n```\nsession_array = test['session'].unique()\nrandom.shuffle(session_array)\nsession_split = np.array_split(session_array, N_FOLDS)\ni = 0\nfor session_subset in session_split:\n     if i == 0:\n          fold_df = pd.DataFrame(zip(session_subset, np.zeros(len(session_subset)) ), columns=['session', 'fold'])\n          fold_df['fold'] = fold_df['fold'].astype(int)\n     else:\n          extra = pd.DataFrame(zip(session_subset, np.zeros(len(session_subset)) ), columns=['session', 'fold'])\n          extra['fold'] = extra['fold'].astype(int)\n          extra['fold'] = extra['fold'] + i\n          fold_df = pd.concat([fold_df, extra], axis=0)\n     i = i+1\nfold_df.to_csv('session_fold.csv', index=False)\n```\n\nThis allowed me to produce the candidate dataframe separately avoiding memory issues. Therefore for each fold I had the whole set of candidates in separate files. With this I was able to perform undersampling to reduce the # of negative samples to 20 times the # of positives samples, and I saved these dataframes for the training. Having them separately made performing n-fold training and computing validation very easily since I had both sampled and original candidates ready.\n\n## Features\nThis was probably the part where I was lacking the most due to time constraint on my part and due to the low score that I got, given an high recall rate on the candidate generation part. I produced features regarding:\n - user\n - item\n - user-item interactions\n - how the candidate was generated\n\nand I saved them as csv files locally, in order to be able to combine them with the required dataframe as I needed.\n\n## Training\nFor training, as I stated before, having the sampled dataframes saved allowed me to combine them easily with the features and also to save them. This made me able to train an XGBoost Ranker locally on my macbook in ~20-30 minutes. I wasn't able to try the LGBM Ranker since it wasn't supported for my current MacOS version (remember to never update to the latest os version haha), while the Catboost ranker has almost the same performances.\n\n## Inference\nFor inference, to avoid out of memory errors I loaded each non-sampled dataframe by chunk, add features to the chunk and then perform inference and finally saving the predictions. At the end I combined the predictions and took the top 20 ones to submit.",
    "2125035": "Congrats! No worries after the leaderboard is finalized you will probably get silver ;)",
    "2125126": "Thank you! I hope so, having that massive drop in ranks and landing outside silver wasn't funny at all...\n\nAlso congrats on your score and on your silver medal!",
    "2125358": ">Word2Vec similar items to the last one seen in the session\nIt is worth trying also min/mean/max to several last aids in session. I guess this alone could give you enough points to stay in silver.",
    "2125454": "This idea was on my to do list, together with trying different co-visitation matrices but unfortunately I didn't have time. \n\nCongrats on your medal!"
  },
  "source": "meta"
}