Interactive Probabilistic Grocery Price Forecasting Demo¶

This notebook demonstrates how to load grocery price datasets, extract features, and use probabilistic forecasts to make smart buying decisions under price uncertainty.

The Core Problem¶

Standard forecasting models output a single predicted price. However, grocery pricing is highly discrete: prices stay flat for weeks, then drop suddenly during sales (promotions). A probabilistic model predicts a mean ($\mu$) and a variance ($\sigma^2$), which allows us to estimate the probability and depth of upcoming promotions to optimize purchase timing.

In [1]:
import os
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from scipy.stats import norm

# Set plots to display inline
%matplotlib inline
sns.set_theme(style="whitegrid")

1. Load Compiled Dataset¶

We load the clean panel dataset containing exactly 13,272 products tracked over 8 weeks.

In [2]:
df = pd.read_csv("../data/grocery_prices.csv")
print(f"Dataset contains {df['upc'].nunique()} unique products across {df['date'].nunique()} weeks.")
print(f"Total rows: {df.shape[0]}")
df.head()
Dataset contains 13272 unique products across 8 weeks.
Total rows: 106176
Out[2]:
date upc name category current_price regular_price is_promotion
0 2026-05-17 4510 Produce Dragon Beans Trash Bags 10.25 10.25 0
1 2026-05-24 4510 Produce Dragon Beans Trash Bags 10.25 10.25 0
2 2026-05-31 4510 Produce Dragon Beans Trash Bags 10.25 10.25 0
3 2026-06-07 4510 Produce Dragon Beans Trash Bags 10.25 10.25 0
4 2026-06-14 4510 Produce Dragon Beans Trash Bags 10.25 10.25 0

2. Inspect Price Volatility¶

Let's see which categories have the highest price volatility (standard deviation over the 8 weeks).

In [3]:
volatility_by_cat = df.groupby("category")["current_price"].std().sort_values(ascending=False).head(10)
plt.figure(figsize=(10, 4))
sns.barplot(x=volatility_by_cat.values, y=volatility_by_cat.index, palette="viridis")
plt.title("Top 10 Grocery Categories by Price Volatility")
plt.xlabel("Price Standard Deviation ($)")
plt.ylabel("Category")
plt.tight_layout()
plt.show()
/var/folders/4x/20x5fv2n5hl244kqtmw_nv0h0000gn/T/ipykernel_58561/3582879436.py:3: FutureWarning: 

Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `y` variable to `hue` and set `legend=False` for the same effect.

  sns.barplot(x=volatility_by_cat.values, y=volatility_by_cat.index, palette="viridis")
No description has been provided for this image

3. Load Engineered Splits & Feature Columns¶

In [4]:
df_train = pd.read_csv("../data/splits/train.csv")
df_test = pd.read_csv("../data/splits/test.csv")

embedding_cols = [c for c in df_train.columns if c.startswith("text_svd_")]
feature_cols = [
    "price_lag_1",
    "promo_lag_1",
    "price_roll_mean_2",
    "price_roll_std_2",
    "price_diff_1",
    "cat_price_mean",
    "cat_promo_rate",
    "regular_price"
] + embedding_cols

print(f"Train set: {df_train.shape[0]} rows")
print(f"Test set: {df_test.shape[0]} rows")
print(f"Features used for training ({len(feature_cols)}):\n", feature_cols)
Train set: 53088 rows
Test set: 13272 rows
Features used for training (18):
 ['price_lag_1', 'promo_lag_1', 'price_roll_mean_2', 'price_roll_std_2', 'price_diff_1', 'cat_price_mean', 'cat_promo_rate', 'regular_price', 'text_svd_0', 'text_svd_1', 'text_svd_2', 'text_svd_3', 'text_svd_4', 'text_svd_5', 'text_svd_6', 'text_svd_7', 'text_svd_8', 'text_svd_9']

4. Probabilistic Buying Decision Engine (Simulation)¶

We implement a decision rule: if the current price of a product is high relative to its predicted distribution's lower bounds, we output a WAIT signal (anticipating a promotion). If the predicted standard deviation is extremely low, it indicates a stable price, outputting a BUY signal.

In [5]:
def make_buying_recommendation(current_price, regular_price, pred_mean, pred_std, threshold_percentile=15):
    """
    Outputs a BUY or WAIT recommendation based on the predicted price distribution.
    """
    # Calculate the target lower percentile of next week's predicted price
    z_score = norm.ppf(threshold_percentile / 100.0)
    percentile_price = pred_mean + z_score * pred_std
    
    # Potential discount depth
    potential_discount = current_price - percentile_price
    discount_ratio = potential_discount / current_price
    
    # Decision logic
    # If next week's 15th percentile is significantly lower than current price (e.g. > 10% lower),
    # it suggests a very high likelihood of a sale. We advise waiting.
    if discount_ratio > 0.10 and pred_std > 0.15:
        recommendation = "WAIT (Promotion Likely Next Week)"
        reason = f"Expected price next week could drop to ${percentile_price:.2f} (a potential {discount_ratio:.1%} savings)."
    else:
        recommendation = "BUY (Price Stable)"
        reason = f"Predicted price is stable near ${pred_mean:.2f} with low volatility ($\\sigma = {pred_std:.2f}$)."
        
    return {
        "Current Price": f"${current_price:.2f}",
        "Regular Price": f"${regular_price:.2f}",
        "Predicted Mean": f"${pred_mean:.2f}",
        "Predicted Std": f"${pred_std:.2f}",
        "15th Percentile": f"${percentile_price:.2f}",
        "Recommendation": recommendation,
        "Explanation": reason
    }

# Demo simulation on sample cases
demo_cases = [
    {"name": "Coca-Cola 12 Pack (Volatile)", "curr": 8.99, "reg": 8.99, "mean": 6.99, "std": 1.50},
    {"name": "Organic Cinnamon Powder (Stable)", "curr": 3.99, "reg": 3.99, "mean": 3.99, "std": 0.01},
    {"name": "Tillamook Cheddar Cheese (On Sale Now)", "curr": 4.99, "reg": 6.49, "mean": 6.20, "std": 0.60}
]

results = []
for case in demo_cases:
    rec = make_buying_recommendation(case["curr"], case["reg"], case["mean"], case["std"])
    rec["Product Name"] = case["name"]
    results.append(rec)
    
pd.DataFrame(results)[["Product Name", "Current Price", "Predicted Mean", "Predicted Std", "Recommendation", "Explanation"]]
Out[5]:
Product Name Current Price Predicted Mean Predicted Std Recommendation Explanation
0 Coca-Cola 12 Pack (Volatile) $8.99 $6.99 $1.50 WAIT (Promotion Likely Next Week) Expected price next week could drop to $5.44 (...
1 Organic Cinnamon Powder (Stable) $3.99 $3.99 $0.01 BUY (Price Stable) Predicted price is stable near $3.99 with low ...
2 Tillamook Cheddar Cheese (On Sale Now) $4.99 $6.20 $0.60 BUY (Price Stable) Predicted price is stable near $6.20 with low ...

5. Review Saved Evaluation Plots¶

Let's look at the generated calibration curve from the training pipeline. If the curves for TensorFlow and PyTorch are close to the diagonal, it shows that the neural networks are well-calibrated and can be trusted to output valid probabilities.

In [6]:
from IPython.display import Image, display

if os.path.exists("reports/plots/calibration_curve.png"):
    print("Loading Calibration Curve:")
    display(Image(filename="reports/plots/calibration_curve.png", width=500))
else:
    print("Calibration plot not found. Make sure to run the main pipeline script first.")
Loading Calibration Curve:
No description has been provided for this image
In [7]:
if os.path.exists("reports/plots/accuracy_comparison.png"):
    print("Loading Model Accuracy Comparison:")
    display(Image(filename="reports/plots/accuracy_comparison.png", width=800))
Loading Model Accuracy Comparison:
No description has been provided for this image